SKN CBBA - ...
SKN CBBA
Cross Border Banking Advisors
SKN | Capital One Builds Evaluation-First Architecture to Govern AI Agents at Scale

Technology

SKN | Capital One Builds Evaluation-First Architecture to Govern AI Agents at Scale

By Or Sushan

•

October 1, 2026

Key Points

  • Capital One has developed an evaluation-first architecture to assess and govern AI agents operating at scale within a regulated banking environment.
  • The framework combines automated scoring, LLM-based evaluation, contextual-accuracy testing and human oversight to identify hallucinations and failures in complex workflows.
  • Capital One’s experience highlights the governance challenges created by autonomous agents, including lost context, contradictory decisions, circular logic and errors that compound across multi-step tasks.

Capital One Treats AI Governance as Infrastructure

Capital One is developing an internal architecture to evaluate and govern AI agents as their use expands across the bank.

Speaking at CoreWeave Fully Connected 2026, Maulin Patel, Managing Vice President of Product, AI and ML Platforms at Capital One, described how the bank is evaluating AI agents in a regulated environment and at the scale required to serve more than 100 million customers.

Patel said the bank cannot govern autonomous AI agents simply by reviewing individual conversations. Instead, Capital One built what he described as an “evaluation first architecture” designed to produce trusted outputs while allowing agentic systems to operate at scale.

The approach places continuous evaluation inside the technology architecture rather than treating governance as a separate process applied after development.

Automated Evaluation Is Combined With Human Oversight

Capital One’s evaluation stack uses several layers to assess AI-agent performance.

The bank tracks scoring metrics and moving averages to build statistical depth around reasoning and workflow performance. It also uses an LLM to evaluate the chain of AI agents, while contextual-accuracy measurements are designed to identify hallucinations.

Human expertise remains part of the process. Capital One supplements and calibrates its automated LLM evaluation channels with subject-matter experts, particularly where complex financial logic is involved.

Patel emphasized that human oversight remains important in regulated industries. Capital One’s subject-matter experts audit complex financial logic while helping calibrate automated evaluation systems.

The structure reflects a hybrid governance model in which automated systems provide continuous monitoring while human reviewers remain involved in areas requiring specialized judgment.

Autonomous Agents Create New Failure Modes

Capital One’s experience also illustrates why evaluating autonomous agents can be more complicated than evaluating a single AI model or conventional chatbot.

Patel identified several problems encountered as agentic systems become more autonomous. In some cases, conversations can lose context over multiple turns, causing the system to forget the original intent. Models can also contradict decisions made earlier in the workflow.

Another issue is circular logic, where an agent repeatedly asks for information that has already been provided. Multi-step tasks create an additional risk because an early mistake can propagate through subsequent actions and compound into a larger failure.

These challenges led Capital One to build its evaluation architecture specifically around autonomous workflows rather than simply extending evaluation methods designed for individual models.

Governance Is Embedded Into the Developer Workflow

Capital One is also positioning governance as part of the development process rather than an external restriction on developers.

Patel said embedding continuous evaluation directly into the developer library allows teams to scale AI systems while maintaining safety controls. In this model, governance becomes part of the development infrastructure and can support faster deployment rather than functioning solely as a compliance checkpoint.

The distinction is significant for financial institutions. AI agents that interact with customer information, financial logic and regulated processes require controls that operate continuously as systems evolve.

For banks, this creates a technology-management challenge alongside the underlying AI opportunity: evaluation must remain effective as agents become more autonomous, workflows become more complex and deployments expand.

AI Infrastructure Becomes Part of Banking Risk Management

Capital One’s broader technology strategy provides the backdrop for its approach to agentic AI. The bank has invested heavily in technology, data and AI infrastructure as part of its effort to integrate acquisitions, improve efficiency and support growth.

Patel also pointed to the importance of its collaboration with CoreWeave and the role of code, infrastructure and human oversight in operating AI-agent evaluations at scale.

For wealth-management and banking organizations, the implications extend beyond generative AI applications. Autonomous systems can increasingly influence customer interactions, financial workflows and internal decision processes, making evaluation architecture an important component of operational and regulatory risk management.

Closing Insights

Capital One’s approach illustrates a shift from simply deploying AI models toward governing networks of autonomous agents. Lost context, contradictory outputs, circular reasoning and compounding errors create risks that become more difficult to manage as agentic systems operate across longer workflows.

The bank’s evaluation-first architecture combines automated measurement with human expertise and embeds continuous testing into development infrastructure. For regulated financial institutions, that model places governance closer to the technology itself, potentially allowing AI deployment to scale while maintaining ongoing oversight.

For a confidential discussion regarding retail banking strategy, insurance distribution models, customer loyalty ecosystems, digital financial services, or cross-border financial innovation opportunities, contact our senior advisory team.

Leave a Reply

Your email address will not be published. Required fields are marked *

More like this

Seraphinite AcceleratorOptimized by Seraphinite Accelerator
Turns on site high speed to be attractive for people and search engines.