What 240 browser-agent runs tell us about choosing a model
Fable led on qualified passes, Flash on cost, and Sol on median task time. Adding Astra and Kimi makes the trade-offs more interesting—not less.
8 models · 6 fixed cases · 5 repetitions each · 240 selected executions · Kernel
Editorial revision: September 9, 2026. Original scores, costs and timings unchanged.
Choosing a model for a browser agent is not the same as choosing the most impressive chatbot. The job may be entering a form, monitoring a construction opportunity or checking the terms of a contract. The useful output is a completed, verifiable workflow—not just an answer that sounds right.
We compared eight models in AI Drive across six browser workflows, with five repetitions per case. The resulting 240 selected executions measure qualified passes, end-to-end task time and settled AI Drive credits. All browser work used Kernel; email delivery was not part of the test.
Four results frame the decision. Claude Fable 5.1 had the highest qualified-pass count, at 29/30. Kimi K3 followed at 28/30. Gemini 3.8 Flash had the lowest complete mean task cost, while GPT-5.6 Sol had the lowest median task time. GPT-6 Astra joined Claude Opus 5 at 27/30. Quality, cost and speed did not point to the same model.
Qualified passes
AI DriveHigher is better · six cases × five repetitions per model
Dollar estimates use 2,000 AI Drive credits per $1, the historical display basis—not provider API prices or a current pricing quote. The results table shows both units.
What we mean by a qualified pass
A run had to meet the frozen rubric: at least 90% field accuracy, every configured critical field correct, the required source evidence, permitted browser behavior, completion within the time limit and an exact settled-credit receipt. One failed requirement is enough to prevent a qualified pass.
That distinction matters. A wrong deadline, an unsupported quotation and an invalid browser action are different failures. We do not count every non-pass as a hallucination. The tasks also ask the agent to distinguish information that is absent from information that is not stated, rather than fill gaps with plausible guesses.
A qualified pass is not a claim of error-free output
Two examples show what this checklist rewards. Fable qualified on all five docket runs while supplying a majority count of 9 where the reviewed answer expected NOT_STATED. The scorer flags this as an unsupported filled-in value, but the field was non-critical, so those runs could still pass. This does not establish that 9 is false; it failed the task's source-only, no-inference requirement.
Astra matched every scored answer field in all 30 runs, but qualified on 27. One quotation failed the frozen-source check, and two docket responses added page-query parameters to the required citation URLs. These are evidence-contract failures, not proof that the scored factual answers were wrong.
The original rubric and scores remain unchanged for every model. The headline measures completion of the full acceptance checklist—not which model made the fewest factual mistakes. The examples do not establish a different universal winner either.
Why runs did not qualify
Counts below cover the 30 attempts per model. A run may fail more than one rule, so the three failure columns can overlap and must not be added together.
Answer-rule failures mean incomplete output, a wrong critical field, or less than 90% field accuracy. Evidence-rule failures mean a required citation or quotation did not pass. Browser / other failures cover execution, safety, provider verification, timing and measurement requirements.
| Claude Fable 5.1 | 1 | 0 | 0 | 1 |
| Kimi K3 | 2 | 0 | 0 | 2 |
| Claude Opus 5 | 3 | 0 | 0 | 3 |
| GPT-6 Astra | 3 | 0 | 3 | 0 |
| GPT-5.6 Sol | 4 | 2 | 4 | 0 |
| Gemini 3.8 Flash | 5 | 0 | 3 | 2 |
| Claude Sonnet 5 | 10 | 1 | 9 | 4 |
| Grok 4.6 | 11 | 7 | 3 | 2 |
This diagnoses non-passes, not every answer mistake. A passing run can still contain non-critical mistakes. Zero in a failure column does not mean zero factual errors; these are not hallucination rates.
The highest score is a starting point, not the whole decision
Fable's 29/30 was the highest observed aggregate. It passed five of the six cases on every repetition, missing one SEC-contract attempt. Kimi and Opus also had five cases with 5/5 results, despite having lower totals. This is why the total and the per-workflow breakdown belong together.
Kimi's 28/30 is especially relevant to the shortlist: its mean was about 2,064 credits per attempted task, versus 7,949 for Fable. That is a useful observed cost-quality trade-off. The one-pass difference between them is too small, over too few distinct cases, to establish a general reliability advantage for either model.
The eight-model chart combines earlier cohorts with a fresh production extension. Collection dates, environments and some execution prompts differ; all use the reviewed scoring packs. These comparisons guide further validation, but they do not isolate model capability from the rest of the system.
Mean cost per attempted task
AI DriveAI Drive creditsEstimated dollars · lower is better
Median task time
AI DriveEnd-to-end elapsed time · lower is better
Cost and time tell different stories
Gemini 3.8 Flash averaged 994 credits per attempted task—about $0.50 at the report's illustrative conversion. It was the lowest-cost model among the eight featured cohorts, with 25/30 qualified passes. That is a sensible candidate for monitored, low-consequence work when it meets the quality requirement of the exact task.
GPT-5.6 Sol had the lowest observed median time: 1 minute 53 seconds. Its mean was 2,441 credits, with 26/30 passes. Flash took 3 minutes 26 seconds at the median. If latency matters, the lowest credit bill need not be the best operational choice.
Astra's median was 2 minutes 6 seconds, but its mean task cost was 8,417 credits—approximately $4.21, the highest in this comparison. It passed 27/30. This batch does not support treating the newer model as an automatic replacement for every workflow; the additional spend needs to be justified on the task being deployed.
Qualified passes vs. task cost
AI DriveHigher and farther left is the favorable direction; each point is one model.
Swipe the chart horizontally to read every label.
Read the workflow, not just the headline
The six cases deliberately mix interaction, monitoring and evidence-heavy research. The form is a public demo using synthetic data. Bid intake and addenda trace use construction notices. The legal-oriented cases use public SEC, eCFR and Supreme Court records. No customer credentials, private bid information or live filing were needed.
Each bar below is five executions of one fixed case. A 5/5 result is a repeatability signal on that case—not five independently sampled customers, contracts or regulations. Use it to choose the next validation task, not to skip validation.
Controlled form
AI DriveFill a multi-step demo with synthetic details, verify the confirmation screen and stop before submitting.
Bid intake
AI DriveExtract an archived construction opportunity and apply a fixed, synthetic qualification filter.
Addenda trace
AI DriveTrace solicitation amendments, preserve old and new deadlines, and identify stated scope changes.
SEC contract
AI DriveExtract contract terms from an SEC exhibit and distinguish missing terms from terms that are actually stated.
eCFR regulation
AI DriveTrace a historical regulation change using supplied point-in-time eCFR records.
Court docket
AI DriveReconcile an official court docket and opinion, including dates, recent events and supporting quotations.
A non-pass does not always mean a wrong answer
Astra's three misses were evidence checks, not failures of the scored factual fields: one quotation did not match the frozen source, and two docket responses added PDF-page query parameters to the required citation URLs. The frozen rubric requires the reviewed URLs exactly. This is a citation-contract failure; it is not proof that the underlying legal answer was false.
Kimi's two misses were browser-execution checks, despite passing its answer and evidence gates. We retained both sets of failures rather than adjusting the rules after seeing which model benefited. The practical lesson is to examine why a run failed, not to reinterpret the qualified-pass rate as a hallucination rate.
Two ways to use these results
High-risk work
When a missed deadline or unsupported contract term carries a large cost, evaluate critical-field errors, unsupported assertions and evidence quality alongside completion rate. Fable achieved the highest end-to-end qualified-pass count in this sample, making it a shortlist candidate—not an established safest model. Astra's complete scored-answer-field matches alongside citation failures show why the failure breakdown matters. Validate the exact workflow and retain human approval before consequential action.
Routine checks
For low-consequence, easily checked work, Flash is the observed economy candidate at about 994 credits per task. Sol is worth comparing when turnaround time matters. Use either only where it clears the task's quality threshold, with monitored samples and a route to human review when evidence is missing.
The takeaway
The useful outcome of a benchmark is a better deployment decision, not a permanent winner. Fable had the highest observed pass count, Flash the lowest mean cost, Sol the lowest median time, and Kimi a strong open-source result. Start from the task, decide what failure would cost, and validate the model at that level.
Methodology and limits
This is a small, controlled benchmark. The limitations below are part of the result.
Scope. Six fixed public-data cases, five repetitions per model, using AI Drive's Kernel browser tools and model catalog defaults. No email trigger. Defaults were not normalized to a common reasoning or token budget. This measures the models as offered in the application, not bare-model capability.
Scoring. The frozen rubric checks field accuracy, critical fields, evidence, permitted execution, time limits and exact credit receipts. Citation matching uses reviewed URLs and quotations. Qualified pass is not a factual-accuracy percentage, and non-pass is not synonymous with hallucination.
Time and cost. Median elapsed time includes unsuccessful attempts. Mean cost is shown only with all 30 receipts; unknown is never zero. Dollar estimates use the historical display basis of 2,000 credits per $1, not a current pricing quote or invoice. Separate maintenance and excluded infrastructure attempts are not added to model task costs.
Collection conditions. The fresh production extension was collected September 6–7, 2026. Scheduling changed from serial to five simultaneous requests during repetition 3, which can affect latency. The deployed backend revision could not be independently attested.
Historical comparison. Six baseline cohorts were collected August 27–September 3 across production and development. Earlier addenda/docket execution packs differ from the corrected packs used for the fresh extension; outputs use the same reviewed scoring packs. The cohorts are not simultaneous or deployment-matched, so differences cannot be attributed to model capability alone.
Open-source selection. Kimi was selected from the five fresh candidates by highest qualified-pass count. The predeclared tiebreaks were complete mean credits, then median time; neither was needed. Selection on these same results is not independent confirmation. Every candidate is retained in the companion article and evidence.
Sample size. Thirty attempts per model repeat just six cases. We make no statistical-significance or universal reliability claim. The cases use supplied URLs, fixed output contracts and explicit missing-value rules; this is not a test of open-ended discovery across the web.
Preservation. The fresh extension has 180 selected cases from 182 physical measured attempts. Two documented infrastructure interruptions received approved replacements; both originals are preserved. Wrong answers, invalid actions and genuine task deadlines were not retried to improve scores. The open-source subset contains 150 selected cases; the broader article includes 180 historical baseline cases plus 60 fresh Astra/Kimi cases.
Reproducibility. The reader-friendly protocol lists the six tasks, time limits, exact AI Drive model labels, recorded-setting limitations, prompt templates, expected outputs and scoring rules. It distinguishes historical execution packs from reviewed scoring packs. No score or run was changed for this editorial update.
Complete results
Select any column heading to sort. Time is the median per attempt; cost is the mean across all 30 selected attempts when all receipts are available. “5/5 workflows” counts cases that qualified on every repetition.
| Claude Fable 5.1 | 29/30 | 5/6 | 7,949.4 | $3.97 | 2m 27s | 30/30 |
| Kimi K3 | 28/30 | 5/6 | 2,064.0 | $1.03 | 3m 11s | 30/30 |
| Claude Opus 5 | 27/30 | 5/6 | 4,936.2 | $2.47 | 2m 24s | 30/30 |
| GPT-6 Astra | 27/30 | 4/6 | 8,417.1 | $4.21 | 2m 06s | 30/30 |
| GPT-5.6 Sol | 26/30 | 3/6 | 2,441.3 | $1.22 | 1m 53s | 30/30 |
| Gemini 3.8 Flash | 25/30 | 3/6 | 993.5 | $0.50 | 3m 26s | 30/30 |
| Claude Sonnet 5 | 20/30 | 2/6 | 2,112.5 | $1.06 | 3m 00s | 30/30 |
| Grok 4.6 | 19/30 | 2/6 | 2,019.4 | $1.01 | 3m 56s | 30/30 |