Job Description
If you are looking for a game-changing career, working for one of the world's leading financial institutions, you’ve come to the right place.
As a Principal Software Engineer -AI & Data Engineering at JPMorgan Chase within the Cloud Foundational Services which is aligned to Infrastructure Platforms, you will be in an extremely hands-on senior engineering role focused on architecting, developing, testing, deploying, operating, debugging, and continuously improving production AI, agentic, data, retrieval, and cloud-native workloads.
Job responsibilities
- Personally design, code, test, deploy, debug, monitor, and improve AI-native services, agent systems, APIs, retrieval services, automation services, data pipelines, and cloud-native components.
- Build reliable agent and multi-agent solutions using Google ADK, MCP, or equivalent frameworks, including secure tool exposure, enterprise context, prompts, resources, APIs, knowledge sources, and reusable capabilities.
- Architects and governs agentic AI-enabled engineering workflows (using enterprise-authorized tools within the work environment) to improve delivery speed, code quality, and operational outcomes at scale (e.g., AI-driven PR review assistance, test generation/maintenance, release readiness checks, incident triage and root-cause acceleration), while defining guardrails for validation, security, resiliency, and reuse across teams.
- Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation at scale.
- Implement tool calling, planner/executor patterns, memory, state management, long-running workflows, human-in-the-loop review, approval gates, access controls, audit trails, guardrails, and safe-failure behavior.
- Develop data ingestion, transformation, curation, metadata enrichment, retrieval pipelines, vector/search pipelines, data quality controls, enterprise connectors, and integration services that support reliable AI grounding.
- Own observability, tracing, telemetry, evaluation, regression testing, prompt and version management, drift monitoring, incident response, runbooks, resiliency testing, performance tuning, and cost optimization.
- Drive design reviews, code reviews, debugging sessions, production incident analysis, technical spikes, reference implementations, and mentoring while remaining accountable for delivery quality.
- Work with product, engineering, platform, risk, security, control, and business stakeholders to convert requirements into practical engineering roadmaps and dependable production systems.
Required qualifications, capabilities, and skills
- Formal training or certification on software engineering concepts and 7+ years applied experience
- Proven experience personally architecting, developing, deploying, operating, debugging, and improving production-grade AI systems, agentic services, APIs, automation services, and cloud-native workloads.
- Hands-on mastery of modern agentic development, with depth in Google Agent Development Kit (ADK) and Model Context Protocol (MCP). Candidates must be able to build and operate production agent systems and expose tools, resources, prompts, context, enterprise knowledge, and external capabilities through secure, reusable interfaces. Candidates without deep ADK and MCP experience must demonstrate equivalent hands-on production depth in LangGraph, CrewAI, OpenAI Agents SDK, AutoGen / Microsoft Agent Framework, Semantic Kernel, LangChain / LlamaIndex agent frameworks, or similar enterprise-grade ecosystems.
- Demonstrated experience designing and leading adoption of agentic AI-enabled development practices (using enterprise-authorized tools within the work environment) across teams, including setting standards for human-in-the-loop validation, auditability/traceability of changes, and secure handling of sensitive data.
- Strong understanding of responsible AI use and control expectations in engineering workflows, including security/resiliency implications, data sensitivity, and risk-based governance; ability to influence senior technical leaders on safe scaling patterns and reuse.
- High proficiency in Python and practical experience with AI and ML engineering frameworks such as LangChain, LlamaIndex, PyTorch, Hugging Face, or equivalent libraries.
- Proven ability to design prompts, tool instructions, few-shot examples, reusable prompt patterns, and chained workflows as versioned engineering assets with testing, review, traceability, regression analysis, and release control.
- Strong experience connecting models and agents to enterprise knowledge using embeddings, vector stores, search, retrieval pipelines, ranking, grounding, context assembly, and evidence-based response generation.
- Experience evaluating proprietary models, open-source models, hosted APIs, self-managed deployments, retrieval strategies, caching, batching, routing, and fallback patterns based on quality, latency, resilience, cost, complexity, and control requirements.
- Practical experience evaluating LLM and agent behavior across task success, factuality, groundedness, policy compliance, hallucination rate, regression behavior, latency, cost, tool-use accuracy, and user experience. Ability to monitor production systems for drift, quality degradation, behavioral regressions, and operational anomalies.
- Experience implementing review, approval, escalation, override, output validation, sensitive-data handling, prompt-injection resistance, secure tool use, decision traceability, and responsible AI controls ; Experience building custom MCP servers, MCP tools, agent tools, connectors, reusable tool registries, and interoperability patterns across multiple agent frameworks.
Preferred qualifications, capabilities, and skills
- Experience with structured architecture models, reusable patterns, and tools or languages such as FINOS CALM or similar frameworks to describe, validate, and automate design decisions.
- Hands-on capability to build, operate, support, and continuously improve an enterprise repository and artifact-store capability for architecture metadata, architecture models, reusable patterns, controls, compliance evidence, and related architecture assets, including assets expressed in standards or formats such as FINOS CALM. Candidates must be able to operate this platform with production discipline, including artifact lifecycle management, access control, traceability, auditability, governance workflows, dependency management, and operational support.
- Experience mitigating prompt injection, tool misuse, data leakage, unauthorized actions, insecure retrieval, data poisoning, and unsafe model behavior.
- Experience with model deployment, model registries, evaluation pipelines, CI/CD for AI services, scalable serving, production AI operations, and model/runtime observability.
J.P. Morgan is a global leader in financial services, providing strategic advice and products to the world’s most prominent corporations, governments, wealthy individuals and institutional investors. Our first-class business in a first-class way approach to serving clients drives everything we do. We strive to build trusted, long-term partnerships to help our clients achieve their business objectives.
We recognize that our people are our strength and the diverse talents they bring to our global workforce are directly linked to our success. We are an equal opportunity employer and place a high value on diversity and inclusion at our company. We do not discriminate on the basis of any protected attribute, including race, religion, color, national origin, gender, sexual orientation, gender identity, gender expression, age, marital or veteran status, pregnancy or disability, or any other basis protected under applicable law. We also make reasonable accommodations for applicants’ and employees’ religious practices and beliefs, as well as mental health or physical disability needs. Visit our FAQs for more information about requesting an accommodation.