Editor’s note: This report is based on our daily use of the models and the testing behind CS Bench. We identify which findings come from the benchmark and which reflect our judgment. The rankings are current as of August 2026.
Exhibit 1 groups the models into five classes based on the work they can complete reliably inside an agent system.
The Diligence Stack has covered the AI market broadly. This is our first full assessment of the models themselves. Something we will keep updating at different points in time. From talking with large enterprises, we believe we are going through the same exercise as it relates to both models that are acceptable for a wide range of knowledge work use cases and costs associated and tokens used. We share our learnings from this exercise and the key takeaways relevant to compute infra and model economics.
We built CS Bench to compare models inside our knowledge base agent that is the basis of our research for the Diligence Stack. The published benchmark measures the quality and estimated cost of first-draft of a financial analysis which is the anchor use case we test. We also use the models every day for research, software work, and the production of finished files. This report uses all tasks tested as its base for analyzing each model.
From a model evaluation standpoint, buyers pay model providers by the token and judge the result by the completed work. A low token price does not help if the output has to be checked for an hour or rebuilt. We include that review time when we compare model cost.
Financial modeling shows the problems more often than other workflows. The most common failure is a workbook that looks finished before its logic has been checked. It may have several tabs and a working scenario switch. The formatting looks credible on a first review. A deeper check then finds revenue pulled from the wrong fiscal period or driver rows hard-coded as values made to look like formulas. Finding those errors can take an hour or more.
We therefore score the work and the presentation separately. Correct work with poor formatting creates cleanup. Work with obvious errors is usually rejected quickly. The costly case is a polished file with errors buried inside it because the presentation makes the work look ready to use.
Exhibit 2 shows the difference. CS Bench measures the first draft before an expert takes over. We are now adding the time needed to review that work, the errors that repeat, and the time needed to find them. Those checks give us a better estimate of the cost of usable work.
Parameter count says little about whether a model will finish the job. We focus on whether it understands an incomplete task, uses tools until the work is done, and recovers when its first approach fails. The software around the model also affects the result. Retrieval, file handling, and execution can make the same model much more or less reliable. Buyers are paying for that full system.
Public API rates describe only one way buyers pay. Small teams often use fixed-price subscriptions. Enterprises negotiate lower rates in exchange for spending commitments that they may not fully use. The real cost includes the model, retries, and human review. A low token price can still lead to a high cost for completed work.
For investors, the amount of review is one way to see which models still earn a premium as token prices fall. Smaller models can handle large volumes of clearly defined work. Larger models remain better suited to work that requires more judgment. Software companies can also earn a share of the value when their products make the models more reliable. The market will support several classes of models because the work and the cost of failure vary widely.
Inside the Full Report
How we group the models and our assessment of each class.
Our scorecard for work quality, presentation, tool use, instruction retention, judgment, and cost.
Workload-level recommendations for research, software engineering, financial modeling, and everyday assistant use.
How subscriptions, enterprise contracts, usage, and review change the cost of completed work.
Why switching models during a long conversation can raise the total cost.
The evidence that would change our rankings and what CS Bench will test next.




