Switching Bluebox to Sonnet 5.5: the eval results
Bluebox runs on Anthropic's Sonnet 5.5 model since October 1. In our evals with SREGym, Sonnet 5.5 diagnosed more problems correctly than Sonnet 5
Engineering insights on production observability, coding agents, and AI-first development. What we build, what we learn, and where the industry is headed.
Bluebox runs on Anthropic's Sonnet 5.5 model since October 1. In our evals with SREGym, Sonnet 5.5 diagnosed more problems correctly than Sonnet 5
You've got a coding agent. It writes code fast. But it has no idea what's happening in production: no error rates, no traces, no
Come meet the Bluebox team, along with 10,000+ other developers, at WeAreDevelopers North America in San Jose this week. September 24th to 25th at the McEnery Convention
Why would you need Bluebox to run an investigation? In theory, your application is always working perfectly in the production environment. While in practice, it is not. One
If you've ever shipped a feature and held your breath waiting to see what broke in production or spent hours re-prompting your coding agent just
Coding agents are genuinely good at understanding a local codebase. Point one at a repository and it can read the architecture, trace the data flow, interpret configuration, and
Bluebox, Kiro, and AWS DevOps Agent bring production intelligence directly into AI-driven software delivery. By connecting development and operations with shared runtime context, teams can build, deploy,
When AI makes implementation faster, context, decision-making, and trust become the real constraints. Here is what we are learning while building Bluebox. This post is part of
Every team shipping with a coding agent eventually faces the same moment. The agent pushes something to production. Something breaks. And someone says: we need better monitoring. The
Routines are now live in production for everyone. Get started with them from your Bluebox sidebar or start from the docs! In this post: We’ll cover what