What 240 browser-agent runs tell us about choosing a model

Fable led on qualified passes, Flash on cost, and Sol on median task time. Adding Astra and Kimi makes the trade-offs more interesting—not less.

8 models · 6 fixed cases · 5 repetitions each · 240 selected executions · Kernel

Editorial revision: September 9, 2026. Original scores, costs and timings unchanged.

Choosing a model for a browser agent is not the same as choosing the most impressive chatbot. The job may be entering a form, monitoring a construction opportunity or checking the terms of a contract. The useful output is a completed, verifiable workflow—not just an answer that sounds right.

We compared eight models in AI Drive across six browser workflows, with five repetitions per case. The resulting 240 selected executions measure qualified passes, end-to-end task time and settled AI Drive credits. All browser work used Kernel; email delivery was not part of the test.

Four results frame the decision. Claude Fable 5.1 had the highest qualified-pass count, at 29/30. Kimi K3 followed at 28/30. Gemini 3.8 Flash had the lowest complete mean task cost, while GPT-5.6 Sol had the lowest median task time. GPT-6 Astra joined Claude Opus 5 at 27/30. Quality, cost and speed did not point to the same model.

Qualified passes

AI Drive

Higher is better · six cases × five repetitions per model

051015202530Claude Fable 5.1: 29/30 qualified passes29/30ClaudeFable 5.1Kimi K3: 28/30 qualified passes28/30KimiK3Claude Opus 5: 27/30 qualified passes27/30ClaudeOpus 5GPT-6 Astra: 27/30 qualified passes27/30GPT-6AstraGPT-5.6 Sol: 26/30 qualified passes26/30GPT-5.6SolGemini 3.8 Flash: 25/30 qualified passes25/30Gemini3.8 FlashClaude Sonnet 5: 20/30 qualified passes20/30ClaudeSonnet 5Grok 4.6: 19/30 qualified passes19/30Grok4.6
Claude Fable 5.1
29/30
Kimi K3
28/30
Claude Opus 5
27/30
GPT-6 Astra
27/30
GPT-5.6 Sol
26/30
Gemini 3.8 Flash
25/30
Claude Sonnet 5
20/30
Grok 4.6
19/30
Every selected attempt counts. A qualified pass combines answer, evidence, execution and measurement requirements; it is not a factual-accuracy percentage.
Chart cost display

Dollar estimates use 2,000 AI Drive credits per $1, the historical display basis—not provider API prices or a current pricing quote. The results table shows both units.

What we mean by a qualified pass

A run had to meet the frozen rubric: at least 90% field accuracy, every configured critical field correct, the required source evidence, permitted browser behavior, completion within the time limit and an exact settled-credit receipt. One failed requirement is enough to prevent a qualified pass.

That distinction matters. A wrong deadline, an unsupported quotation and an invalid browser action are different failures. We do not count every non-pass as a hallucination. The tasks also ask the agent to distinguish information that is absent from information that is not stated, rather than fill gaps with plausible guesses.

A qualified pass is not a claim of error-free output

Two examples show what this checklist rewards. Fable qualified on all five docket runs while supplying a majority count of 9 where the reviewed answer expected NOT_STATED. The scorer flags this as an unsupported filled-in value, but the field was non-critical, so those runs could still pass. This does not establish that 9 is false; it failed the task's source-only, no-inference requirement.

Astra matched every scored answer field in all 30 runs, but qualified on 27. One quotation failed the frozen-source check, and two docket responses added page-query parameters to the required citation URLs. These are evidence-contract failures, not proof that the scored factual answers were wrong.

The original rubric and scores remain unchanged for every model. The headline measures completion of the full acceptance checklist—not which model made the fewest factual mistakes. The examples do not establish a different universal winner either.

Why runs did not qualify

Counts below cover the 30 attempts per model. A run may fail more than one rule, so the three failure columns can overlap and must not be added together.

Answer-rule failures mean incomplete output, a wrong critical field, or less than 90% field accuracy. Evidence-rule failures mean a required citation or quotation did not pass. Browser / other failures cover execution, safety, provider verification, timing and measurement requirements.

Claude Fable 5.11001
Kimi K32002
Claude Opus 53003
GPT-6 Astra3030
GPT-5.6 Sol4240
Gemini 3.8 Flash5032
Claude Sonnet 510194
Grok 4.611732

This diagnoses non-passes, not every answer mistake. A passing run can still contain non-critical mistakes. Zero in a failure column does not mean zero factual errors; these are not hallucination rates.

The highest score is a starting point, not the whole decision

Fable's 29/30 was the highest observed aggregate. It passed five of the six cases on every repetition, missing one SEC-contract attempt. Kimi and Opus also had five cases with 5/5 results, despite having lower totals. This is why the total and the per-workflow breakdown belong together.

Kimi's 28/30 is especially relevant to the shortlist: its mean was about 2,064 credits per attempted task, versus 7,949 for Fable. That is a useful observed cost-quality trade-off. The one-pass difference between them is too small, over too few distinct cases, to establish a general reliability advantage for either model.

The eight-model chart combines earlier cohorts with a fresh production extension. Collection dates, environments and some execution prompts differ; all use the reviewed scoring packs. These comparisons guide further validation, but they do not isolate model capability from the rest of the system.

Mean cost per attempted task

AI Drive

AI Drive creditsEstimated dollars · lower is better

Gemini 3.8 Flash
994$0.50
Grok 4.6
2,019$1.01
Kimi K3
2,064$1.03
Claude Sonnet 5
2,112$1.06
GPT-5.6 Sol
2,441$1.22
Claude Opus 5
4,936$2.47
Claude Fable 5.1
7,949$3.97
GPT-6 Astra
8,417$4.21
All eight featured cohorts have 30/30 settled receipts. Cost includes selected successful and unsuccessful attempts.

Median task time

AI Drive

End-to-end elapsed time · lower is better

GPT-5.6 Sol
1m 53s
GPT-6 Astra
2m 06s
Claude Opus 5
2m 24s
Claude Fable 5.1
2m 27s
Claude Sonnet 5
3m 00s
Kimi K3
3m 11s
Gemini 3.8 Flash
3m 26s
Grok 4.6
3m 56s
Includes unsuccessful attempts and captured deadlines. This is not tokens per second or time per successful task.

Cost and time tell different stories

Gemini 3.8 Flash averaged 994 credits per attempted task—about $0.50 at the report's illustrative conversion. It was the lowest-cost model among the eight featured cohorts, with 25/30 qualified passes. That is a sensible candidate for monitored, low-consequence work when it meets the quality requirement of the exact task.

GPT-5.6 Sol had the lowest observed median time: 1 minute 53 seconds. Its mean was 2,441 credits, with 26/30 passes. Flash took 3 minutes 26 seconds at the median. If latency matters, the lowest credit bill need not be the best operational choice.

Astra's median was 2 minutes 6 seconds, but its mean task cost was 8,417 credits—approximately $4.21, the highest in this comparison. It passed 27/30. This batch does not support treating the newer model as an automatic replacement for every workflow; the additional spend needs to be justified on the task being deployed.

Qualified passes vs. task cost

AI Drive

Higher and farther left is the favorable direction; each point is one model.

More passes, lower cost0510152025300$0.002,500$1.255,000$2.507,500$3.7510,000$5.00Claude Fable 5.1: 29/30; 7,949.4 credits / $3.97Claude Fable 5.1Kimi K3: 28/30; 2,064.0 credits / $1.03Kimi K3Claude Opus 5: 27/30; 4,936.2 credits / $2.47Claude Opus 5GPT-6 Astra: 27/30; 8,417.1 credits / $4.21GPT-6 AstraGPT-5.6 Sol: 26/30; 2,441.3 credits / $1.22GPT-5.6 SolGemini 3.8 Flash: 25/30; 993.5 credits / $0.50Gemini 3.8 FlashClaude Sonnet 5: 20/30; 2,112.5 credits / $1.06Claude Sonnet 5Grok 4.6: 19/30; 2,019.4 credits / $1.01Grok 4.6Mean AI Drive credits per attempted taskEstimated dollars per attempted task · lower is betterQualified passes / 30

Swipe the chart horizontally to read every label.

Shading indicates direction only, not a validated deployment threshold. Five repetitions of six fixed cases do not establish real-world reliability.

Read the workflow, not just the headline

The six cases deliberately mix interaction, monitoring and evidence-heavy research. The form is a public demo using synthetic data. Bid intake and addenda trace use construction notices. The legal-oriented cases use public SEC, eCFR and Supreme Court records. No customer credentials, private bid information or live filing were needed.

Each bar below is five executions of one fixed case. A 5/5 result is a repeatability signal on that case—not five independently sampled customers, contracts or regulations. Use it to choose the next validation task, not to skip validation.

Controlled form

AI Drive

Fill a multi-step demo with synthetic details, verify the confirmation screen and stop before submitting.

Claude Fable 5.1
5/5
Claude Opus 5
5/5
Claude Sonnet 5
5/5
GPT-5.6 Sol
5/5
GPT-6 Astra
5/5
Grok 4.6
5/5
Kimi K3
5/5
Gemini 3.8 Flash
4/5
Qualified passes / 5; same case repeated five times.

Bid intake

AI Drive

Extract an archived construction opportunity and apply a fixed, synthetic qualification filter.

Claude Fable 5.1
5/5
Claude Opus 5
5/5
Claude Sonnet 5
5/5
GPT-5.6 Sol
5/5
GPT-6 Astra
5/5
Gemini 3.8 Flash
5/5
Kimi K3
5/5
Grok 4.6
4/5
Qualified passes / 5; same case repeated five times.

Addenda trace

AI Drive

Trace solicitation amendments, preserve old and new deadlines, and identify stated scope changes.

Claude Fable 5.1
5/5
Claude Opus 5
5/5
GPT-6 Astra
5/5
Grok 4.6
5/5
Kimi K3
5/5
GPT-5.6 Sol
4/5
Claude Sonnet 5
3/5
Gemini 3.8 Flash
2/5
Qualified passes / 5; same case repeated five times.

SEC contract

AI Drive

Extract contract terms from an SEC exhibit and distinguish missing terms from terms that are actually stated.

Kimi K3
5/5
Claude Fable 5.1
4/5
GPT-6 Astra
4/5
Gemini 3.8 Flash
4/5
GPT-5.6 Sol
3/5
Claude Opus 5
2/5
Grok 4.6
1/5
Claude Sonnet 5
0/5
Qualified passes / 5; same case repeated five times.

eCFR regulation

AI Drive

Trace a historical regulation change using supplied point-in-time eCFR records.

Claude Fable 5.1
5/5
Claude Opus 5
5/5
GPT-5.6 Sol
5/5
GPT-6 Astra
5/5
Gemini 3.8 Flash
5/5
Kimi K3
5/5
Claude Sonnet 5
4/5
Grok 4.6
0/5
Qualified passes / 5; same case repeated five times.

Court docket

AI Drive

Reconcile an official court docket and opinion, including dates, recent events and supporting quotations.

Claude Fable 5.1
5/5
Claude Opus 5
5/5
Gemini 3.8 Flash
5/5
GPT-5.6 Sol
4/5
Grok 4.6
4/5
Claude Sonnet 5
3/5
GPT-6 Astra
3/5
Kimi K3
3/5
Qualified passes / 5; same case repeated five times.

A non-pass does not always mean a wrong answer

Astra's three misses were evidence checks, not failures of the scored factual fields: one quotation did not match the frozen source, and two docket responses added PDF-page query parameters to the required citation URLs. The frozen rubric requires the reviewed URLs exactly. This is a citation-contract failure; it is not proof that the underlying legal answer was false.

Kimi's two misses were browser-execution checks, despite passing its answer and evidence gates. We retained both sets of failures rather than adjusting the rules after seeing which model benefited. The practical lesson is to examine why a run failed, not to reinterpret the qualified-pass rate as a hallucination rate.

Two ways to use these results

High-risk work

When a missed deadline or unsupported contract term carries a large cost, evaluate critical-field errors, unsupported assertions and evidence quality alongside completion rate. Fable achieved the highest end-to-end qualified-pass count in this sample, making it a shortlist candidate—not an established safest model. Astra's complete scored-answer-field matches alongside citation failures show why the failure breakdown matters. Validate the exact workflow and retain human approval before consequential action.

Routine checks

For low-consequence, easily checked work, Flash is the observed economy candidate at about 994 credits per task. Sol is worth comparing when turnaround time matters. Use either only where it clears the task's quality threshold, with monitored samples and a route to human review when evidence is missing.

The takeaway

The useful outcome of a benchmark is a better deployment decision, not a permanent winner. Fable had the highest observed pass count, Flash the lowest mean cost, Sol the lowest median time, and Kimi a strong open-source result. Start from the task, decide what failure would cost, and validate the model at that level.

Methodology and limits

This is a small, controlled benchmark. The limitations below are part of the result.

Scope. Six fixed public-data cases, five repetitions per model, using AI Drive's Kernel browser tools and model catalog defaults. No email trigger. Defaults were not normalized to a common reasoning or token budget. This measures the models as offered in the application, not bare-model capability.

Scoring. The frozen rubric checks field accuracy, critical fields, evidence, permitted execution, time limits and exact credit receipts. Citation matching uses reviewed URLs and quotations. Qualified pass is not a factual-accuracy percentage, and non-pass is not synonymous with hallucination.

Time and cost. Median elapsed time includes unsuccessful attempts. Mean cost is shown only with all 30 receipts; unknown is never zero. Dollar estimates use the historical display basis of 2,000 credits per $1, not a current pricing quote or invoice. Separate maintenance and excluded infrastructure attempts are not added to model task costs.

Collection conditions. The fresh production extension was collected September 6–7, 2026. Scheduling changed from serial to five simultaneous requests during repetition 3, which can affect latency. The deployed backend revision could not be independently attested.

Historical comparison. Six baseline cohorts were collected August 27–September 3 across production and development. Earlier addenda/docket execution packs differ from the corrected packs used for the fresh extension; outputs use the same reviewed scoring packs. The cohorts are not simultaneous or deployment-matched, so differences cannot be attributed to model capability alone.

Open-source selection. Kimi was selected from the five fresh candidates by highest qualified-pass count. The predeclared tiebreaks were complete mean credits, then median time; neither was needed. Selection on these same results is not independent confirmation. Every candidate is retained in the companion article and evidence.

Sample size. Thirty attempts per model repeat just six cases. We make no statistical-significance or universal reliability claim. The cases use supplied URLs, fixed output contracts and explicit missing-value rules; this is not a test of open-ended discovery across the web.

Preservation. The fresh extension has 180 selected cases from 182 physical measured attempts. Two documented infrastructure interruptions received approved replacements; both originals are preserved. Wrong answers, invalid actions and genuine task deadlines were not retried to improve scores. The open-source subset contains 150 selected cases; the broader article includes 180 historical baseline cases plus 60 fresh Astra/Kimi cases.

Reproducibility. The reader-friendly protocol lists the six tasks, time limits, exact AI Drive model labels, recorded-setting limitations, prompt templates, expected outputs and scoring rules. It distinguishes historical execution packs from reviewed scoring packs. No score or run was changed for this editorial update.

Complete results

Select any column heading to sort. Time is the median per attempt; cost is the mean across all 30 selected attempts when all receipts are available. “5/5 workflows” counts cases that qualified on every repetition.

Claude Fable 5.129/305/67,949.4$3.972m 27s30/30
Kimi K328/305/62,064.0$1.033m 11s30/30
Claude Opus 527/305/64,936.2$2.472m 24s30/30
GPT-6 Astra27/304/68,417.1$4.212m 06s30/30
GPT-5.6 Sol26/303/62,441.3$1.221m 53s30/30
Gemini 3.8 Flash25/303/6993.5$0.503m 26s30/30
Claude Sonnet 520/302/62,112.5$1.063m 00s30/30
Grok 4.619/302/62,019.4$1.013m 56s30/30

Original AI Drive measurements and analysis. The chart vocabulary is inspired by Artificial Analysis and its benchmark articles. We are not affiliated with Artificial Analysis; these scores and task costs are not its indices or API-price measurements.