OpenTelemetry and Bluebox: Production Visibility in One Prompt
You've got a coding agent. It writes code fast. But it has no idea what's happening in production: no error rates, no traces, no
Bluebox runs on Anthropic's Sonnet 5.5 model since October 1. In our evals with SREGym, Sonnet 5.5 diagnosed more problems correctly than Sonnet 5 (86.7% vs 73.3%), used 3.3x fewer tokens, finished 3.7x faster, and cost 4.3x less per investigation.
We evaluate Bluebox continuously, against our own benchmarks and against public ones. SREGym is a public, open benchmark of fault scenarios. Each scenario has a target application on Kubernetes and an injected fault, with a known ground truth and an LLM-as-a-judge that scores the investigation.
We ran 20 SREGym problems, 3 attempts each. We ran the suite once with Sonnet 5 and once with Sonnet 5.5. Nothing else changed: the same graders, ground truth, environment, and 1500-second timeout.
| Sonnet 5 | Sonnet 5.5 | Change | |
|---|---|---|---|
| Correct diagnoses | 73.3% | 86.7% | +13.4 points |
| Tokens | ~5.57M | ~1.69M | 3.3x fewer |
| Time* | 607 s | 162 s | 3.7x faster |
| Cost | $2.56 | $0.60 | 4.3x less |
*Agent execution time. It excludes pod startup.
Sonnet 5 made about twice as many model calls per attempt (47.5 vs 20.9), and each call carried more tokens. Sonnet 5.5 got to an answer with less work.
Investigations in Bluebox are now faster and more accurate. Ask Bluebox to investigate an incident and you'll typically get a diagnosis in under three minutes instead of ten, with more correct root causes along the way. It also costs far less to run, so you can ask more questions and investigate more incidents on the same budget.
To note, we cannot compare these results with the public SREGym leaderboard. We run a private fork of SREGym with changes:
The results are valid for one comparison: Sonnet 5 vs Sonnet 5.5 on Bluebox, but we want to keep in mind that this comparison is early:
We'll re-run the same 20 problems with richer Kubernetes data. This brings our setup closer to the public SREGym leaderboard, though our results still won't be directly comparable. We will cover the full methodology in a later post.
You've got a coding agent. It writes code fast. But it has no idea what's happening in production: no error rates, no traces, no
Come meet the Bluebox team, along with 10,000+ other developers, at WeAreDevelopers North America in San Jose this week. September 24th to 25th at the McEnery Convention
Why would you need Bluebox to run an investigation? In theory, your application is always working perfectly in the production environment. While in practice, it is not. One