Skip to content

Switching Bluebox to Sonnet 5.5: the eval results

Switching Bluebox to Sonnet 5.5: the eval results

Bluebox runs on Anthropic's Sonnet 5.5 model since October 1. In our evals with SREGym, Sonnet 5.5 diagnosed more problems correctly than Sonnet 5 (86.7% vs 73.3%), used 3.3x fewer tokens, finished 3.7x faster, and cost 4.3x less per investigation.

Why SREGym

We evaluate Bluebox continuously, against our own benchmarks and against public ones. SREGym is a public, open benchmark of fault scenarios. Each scenario has a target application on Kubernetes and an injected fault, with a known ground truth and an LLM-as-a-judge that scores the investigation.

We ran 20 SREGym problems, 3 attempts each. We ran the suite once with Sonnet 5 and once with Sonnet 5.5. Nothing else changed: the same graders, ground truth, environment, and 1500-second timeout.

Results

Sonnet 5 Sonnet 5.5 Change
Correct diagnoses 73.3% 86.7% +13.4 points
Tokens ~5.57M ~1.69M 3.3x fewer
Time* 607 s 162 s 3.7x faster
Cost $2.56 $0.60 4.3x less

*Agent execution time. It excludes pod startup.

Sonnet 5 made about twice as many model calls per attempt (47.5 vs 20.9), and each call carried more tokens. Sonnet 5.5 got to an answer with less work.

What this means for Bluebox users

Investigations in Bluebox are now faster and more accurate. Ask Bluebox to investigate an incident and you'll typically get a diagnosis in under three minutes instead of ten, with more correct root causes along the way. It also costs far less to run, so you can ask more questions and investigate more incidents on the same budget.

Limits of this approach

To note, we cannot compare these results with the public SREGym leaderboard. We run a private fork of SREGym with changes:

  • The bluebox agent has no direct access to kubectl. Agents on the public leaderboard do.
  • We extended and altered the telemetry, and we use Dynatrace as the observability backend. The public benchmark uses open-source tooling.

The results are valid for one comparison: Sonnet 5 vs Sonnet 5.5 on Bluebox, but we want to keep in mind that this comparison is early:

  • One run per model, on 20 problems with 3 attempts each. The attempts at one problem are not independent, and we have not measured run-to-run variance.
  • The 20 problems ran with OpenTelemetry data only, so some scenarios that need Kubernetes data are harder than they would be with full access.
  • The accuracy gain is 8 more passes out of 60. We treat it as a signal. The gains in tokens, time, and cost are much larger.

Next steps

We'll re-run the same 20 problems with richer Kubernetes data. This brings our setup closer to the public SREGym leaderboard, though our results still won't be directly comparable. We will cover the full methodology in a later post.