> ## Content Index
> Fetch the complete content index at: https://blog.bluebox.ai/llms.txt
> Use this file to discover other available public pages before exploring further.

# OpenTelemetry and Bluebox: Production Visibility in One Prompt
- URL: https://blog.bluebox.ai/opentelemetry-and-bluebox-production-visibility-in-one-prompt/
- Published: 2026-09-28T16:27:55.000Z
- Updated: 2026-09-28T16:27:55.000Z
- Author: Jason Ostroski

---

You've got a coding agent. It writes code fast. But it has no idea what's happening in production: no error rates, no traces, no idea which endpoints are quietly dying under load. 

That's exactly the gap Bluebox is built to close as the industry’s first end-to-end Observability Agent, and OpenTelemetry is the data that fuels Bluebox.

---

## What is OpenTelemetry and Why does Bluebox Need It?

[OpenTelemetry](https://opentelemetry.io/?ref=blog.bluebox.ai) (OTel) is the open standard for collecting telemetry from your running services: traces, metrics, and logs.   
  
OTel is a CNCF project (the same organization behind Kubernetes) and is truly vendor-neutral with stable SDKs and APIs for data collection across every major language and one transfer format (OTLP) accepted by virtually every observability backend. That means you instrument once and can point it wherever you'd like, whether that's Bluebox, Dynatrace, Grafana, or anything else, without re-instrumenting. Your data is never tied to your tooling choices, and your existing OTel investment carries over directly. And because it supports auto-instrumentation across popular frameworks, you're not manually wrapping every function call.  
  
Of the three telemetry types, traces are the most powerful for troubleshooting applications. A **trace** follows a single request through your entire service topology, from frontend web servers, through all the downstream services, to databases and back. Each step is a span, providing all of the deep-dive details and timings about that step. Together, traces and spans give you the full picture: not just that something failed, but exactly where, how long it took, and what everything around it was doing at the time.

One of Bluebox's core loops is: **production signal → root cause → evidence-backed fix**. When something goes wrong (elevated error rates, latency spikes, unexpected service behavior), Bluebox automatically detects issues, kicks off an investigation, pinpoints the precise root cause, and creates an evidence-backed remediation plan attached to your GitHub repo. The evidence and telemetry it uses for all of this is OTel, and traces specifically are how Bluebox builds its topology map and understands how your services call each other.

Your coding agent benefits too. With OTel flowing, the agent can query Bluebox directly whenever it needs production context, whether it's fixing an issue or building a new feature:

```
"What are the most common errors in the last 7 days?"
"Which endpoints in the payments service have the highest p99 latency?"
"Is the checkout flow behaving normally compared to last week?"
"What changed in the last deploy that could explain this error spike?"
```

Instead of reasoning from code alone, the agent is working from real traffic patterns, real error rates, and real service dependencies. It's writing to the system as it *actually* behaves.

---

## Instrumenting with Bluebox: One Prompt Does It All

Bluebox takes all of the guesswork out of installing OpenTelemetry. After you install the Bluebox CLI, you paste one prompt into your coding agent:

```
Add OpenTelemetry to this project — use the Bluebox instrumentation skill
```

  
The agent loads the Bluebox instrumentation skill and starts by inventorying your service: language, framework, what telemetry you already have (if any), and whether containers are involved. 

![](https://blog.bluebox.ai/content/images/2026/09/Screenshot-2026-09-22-at-4.03.36---PM.png)

Discovery and inventory process.

It then builds a full instrumentation plan: which packages to install (with versions resolved live from the registry), which files to create or modify, which recipe to use for your language and framework, and which environment name to tag your telemetry with. You see the plan, including an estimated time to complete, before anything is written.

![](https://blog.bluebox.ai/content/images/2026/09/Screenshot-2026-09-22-at-4.02.25---PM.png)

Bluebox and Kiro's Plan for instrumentation.

Before it writes a single file, it asks you one question: **what level of instrumentation do you want?**

- **Level 1** — traces only
- **Level 2** — traces + metrics
- **Level 3** — traces + metrics + logs, with your existing logger bridged

**Pick your level and it gets to work!**

The Bluebox CLI fetches the right OTLP endpoint for your workspace automatically via `bluebox otlp-endpoint`. The one thing the agent won't touch is your ingest token, it stops and asks you to copy it from the Setup page in your browser yourself. It never appears in any file or CLI output, keeping your token safe by design.

Once wiring is done, the agent starts your service, drives test traffic, and verifies telemetry is flowing with `bluebox ask` — checking that traces, metrics, and logs are all arriving before it declares the job complete. If something is missing, it catches it here. 

![](https://blog.bluebox.ai/content/images/2026/09/Screenshot-2026-09-22-at-4.12.16---PM-2.png)

Bluebox and Kiro validated Otel instrumentation, found an error a logging SDK, fixed it and verified all without human intervention.

In our own instrumentation run, traces and metrics arrived immediately, but `bluebox ask` reported zero log records. The agent dug into the OTel SDK source, found the bug (a constructor API change between versions that caused log exports to silently fail), fixed it, and re-verified, all without having to manually diagnose it.

---

## Already Have OTel? Just Point It at Bluebox

If your services are already instrumented, you're most of the way there. Head to the Bluebox Setup page, where you'll find your OTLP endpoint URL and ingest token ready to copy. Set the URL and token and redeploy your app. That's it, your existing OTel instrumentation now feeds Bluebox. 

---

##   
Why OTel Instrumentation Is Harder Than It Looks

OpenTelemetry is very powerful, but it can be a bit tricky to get right. OTel has to be initialized before any other libraries load or their spans are silently dropped. You need to choose between auto and manual instrumentation. You need to get the OTLP endpoint format, auth header, and token management right. Context propagation breaks across async boundaries and message queues. 

![](https://blog.bluebox.ai/content/images/2026/09/image.png)

Sometime instrumentation feels like: 1\. Set OTEL\_SERVICE\_NAME 2.Configure everything else

Any misconfiguration along the way means missing data, and every failure looks the same: no data arrives, and you can't tell if it's your code, the network, the backend, or a token issue. That's why we made it a priority to make OTel instrumentation across all your services super easy with Bluebox.

---

## What You Get Once It's Flowing

Once telemetry arrives, Bluebox works in the background continuously, monitoring your services, detecting elevated error rates and latency spikes, and surfacing **Findings** on your Overview page. Select **Investigate** on any finding and Bluebox traces it back to root cause, builds an evidence-backed report, and files a GitHub issue with a proposed fix your coding agent can act on immediately.

The loop you're building: ship fast with your coding agent → catch what it missed in production → trace it back → fix it. All without leaving your terminal or coding agent.

---

*Get started at* [*app.bluebox.ai*](https://app.bluebox.ai/?ref=blog.bluebox.ai) *or read the* [*full setup guide*](https://docs.bluebox.ai/getting-started?ref=blog.bluebox.ai)*.*