The AI model you choose is the least important decision you’ll make

I tested four LLMs on the work a capital-markets institution actually performs. The implication is not to buy cheaper AI. It is to invest where advantage is actually created. 

By Prerit Ahuja, Director, Global AI & Data Strategy, DNB Carnegie 

 

I tested more than 100 pieces of model-generated work across five capital-markets workflows. The result challenged a common assumption behind enterprise AI spending: among leading models, paying more did not produce more reliable financial work.

 

For C-suite leaders, there are two consequences. First, do not let model price stand in for business value: among leading LLMs a larger bill may buy little improvement. Second, redirect more of the AI budget towards six priorities: data, workflow design, controls, integration, evaluation, and people. That is how a model available to every competitor becomes an advantage specific to your business. 

 

The objective is not to deploy the best model. It should be to build the best-performing business system. 

Experiment methodology

 

The evaluation covered five representative workflows: extracting bond terms, analysing equity terms, retrieving precedent, synthesising issuer profiles, and assessing M&A comparables. I used a balanced mix of real regulatory filings and transaction data, alongside synthetic content designed to support statistically consistent and repeatable testing. Each model received the same task, source material, and required output.

 

Then, I assessed the completed work using the Capital Markets LLM Reliability Score, or CM-LRS, a seven part framework I developed for high-stakes financial workflows. Four independent AI evaluators from three model families scored the same work, reducing the risk of one model favouring its own answers.

 

The result was consistent: among the three leading commercial models tested, paying more did not produce meaningfully greater reliability. A less costly model matched—and marginally outscored—the most expensive model tested, at around 40% of the cost per representative call. 

Untitled design-Sep-14-2026-09-40-20-3353-AM

 

Same reliability. Opus costs ~1.7x - about $105,000 a year more, for no measurable gain. 

Key assumption: approximately 2 bn aggregate tokens per month across 10 employees, assuming  input (70%) and output (30%) token mix to both models. 

 

This is not a general claim that cheaper is always better. It is a more useful decision rule: for financial services workflows, price and reliability do not move together. When the leading models all meet the required standard, model choice contributes less to competitive advantage than the system built around it. 

 

“Choose a model sufficiently capable for the task. Invest the difference in the capability around it.” 

 

Lesson 1: leading models were effectively tied 

Public LLM benchmarks typically rank general capabilities such as reasoning, coding, or question answering. Financial institutions buy something different: a completed piece of work that must be repeatedly reliable, accurate, sourced, numerically consistent, and ready to support a decision. 

 

CM-LRS tests the qualities that make financial work dependable, from evidence and numerical consistency to completeness and decision usefulness. Across the five workflows tested, the reliability scores of the leading commercial models were so close that none showed a consistent relative advantage.

 

That changes the procurement question. Instead of asking, “Which model leads the public LLM benchmark?”, ask, “What standard must this workflow meet, and which model meets it at the lowest total cost?”

 

The freely available open-source model was cheaper and faster, but weaker overall, particularly when retrieving and synthesising evidence across documents. The strategic implication is not to default to the cheapest model. It is to define the standard first, then avoid paying for model capability the workflow does not need. 

Lesson 2: workflow design mattered more than model prestige 

An LLM never operates alone in a regulated institution. It sits inside a system of data, permissions, prompts, calculations, output formats, controls, and human review. That surrounding design determines much, including whether these inter-connected assets are viewed and actioned upon holistically.

 

We saw this in the results. The less costly model produced concise, disciplined answers that suited structured financial tasks. The more expensive model often reasoned more elaborately, but occasionally omitted requested fields and lost points on workflow completeness. More model capability did not compensate for a weaker fit between the response and the task.

 

This is where institutions should redirect attention and budget. Six investments matter: 

 

  • Data: make proprietary information searchable, permissioned and traceable, with clear ownership.
  • Workflow: specify the task, required fields, sources, output and escalation points before selecting the model.
  • Controls: place calculations and fixed rules in conventional software (deterministic), then test outputs against a defined reliability threshold.
  • Integration: connect the model to the systems where work begins and decisions are recorded, rather than creating another standalone tool.
  • Evaluation: monitor accuracy, completeness, review time and exceptions by workflow, with clear ownership when performance falls below standard.
  • People: assign accountable experts to material decisions and give reviewers the sources, calculations and exceptions needed for a quick decision. 

“Predict with the model, calculate with code, decide with the human.”

 

Use the LLM where language, context, and ambiguity matter; use conventional software for arithmetic and controls; keep an accountable person on material decisions. 

 

Lesson 3: review effort decides the economics 

Model price is only one component of cost. In financial services, the larger cost can be the time required for a highly paid professional to verify, repair, or reconstruct the output.

 

A model can look accurate and still be commercially useless. If the reviewer must find every source, rebuild every calculation, and restore every omitted step, the system has simply created additional review work. That is why decision usefulness—whether a professional could act without materially reworking the output—was the clearest practical indicator of production readiness in the study.

 

The board-level measure should therefore be the total cost of completed, reviewable work. Boards should ask what proportion of outputs clears review without material correction and how much expert time is released. They should then test whether the institution is making decisions faster, serving more clients, reducing risk, or growing without costs rising at the same rate.

 

Over many years, I have observed, redesigned, and built AI-first financial workflows across equity capital markets, debt capital markets, M&A, and adjacent advisory businesses. The recurring lesson is that value does not come from placing a general-purpose model beside an existing process. It comes from rethinking how information is retrieved, analysed, checked, and converted into a decision.

 

That is why the relevant outcome is operating leverage: can the institution handle more work, make better decisions and serve clients faster without cost and risk rising at the same rate? No model can deliver that in isolation. It requires a broader system of data, workflow design, controls and management discipline. 

 

Closing thoughts

In a controlled study of real capital-markets work, a model costing around 40% less matched the reliability of the most expensive model tested.

 

The leading models will continue to change places; prices will fall. Today’s advanced capabilities will quickly become widely available. An institution that builds its AI strategy around access to a particular model is building on a temporary advantage.

 

The durable advantage is knowing where AI creates value, using no more model than the work requires, and investing the difference in workflows that help people make better decisions. The model matters, yes, but do not mistake it for the strategy. 

 

Supporting scientific research published here: arXiv and SSRN.

 

Insight basis: The CM-LRS analysis tested Claude Sonnet 4.6, Claude Opus 4.7, GPT-5.5 and Llama 3.3 70B across five capital-markets workflow classes. Four-evaluator average scores for the three leading commercial models were 4.31, 4.30 and 4.09 out of five. Indicative May 2026 list-price estimates for a representative 5,000-token input and 2,000-token output call were $0.04–$0.06 for Sonnet and $0.20–$0.30 for Opus. 

 

Mask group-2

SUBMIT A COMMENT

We love getting input from our communities, please feel free to share your thoughts on this article. Simply leave a comment below and one of our moderators will review
Mask group

Join the community

To join the HotTopics Community and gain access to our exclusive content, events and networking opportunities simply fill in the form below.

Mask group