TL;DR
The pager moved. My old Nagios alert was “service down.” My new alert is “groundedness score dropped from 0.91 to 0.74 over the last 1,000 requests, and we’re spending $3.20 per resolved query.” Same job. Different physics. AI workloads broke the deterministic model in ways the cloud never did — there is no exit code 0 for “this answer was correct.” The new SLO is cost-to-resolution. This post is about the observability stack that gets you there: OpenTelemetry’s GenAI semantic conventions, OpenLIT for drop-in instrumentation, LLM-as-a-Judge for quality scoring, and why the architecture you already built for Parts 1 through 4 is mostly the right pipe.
The alert that doesn’t fit any old framework
A few months ago, an agentic system I was helping debug started returning answers that were technically formatted correctly, polite, well-structured — and quietly wrong. Not 500 errors. Not timeouts. The model was confidently citing documents that didn’t exist.
There was nothing in any traditional monitoring stack that could have caught this. CPU was fine. Latency was fine. Error rate was zero. The 200 OKs were the bug.
This is Part 5 of The Probability Pivot, and it’s the one where the deterministic vocabulary finally runs out. We’ve been working toward this since Part 1’s exit codes. The arc is complete: we no longer monitor whether something ran. We monitor whether something worked, was true, and was worth what it cost.
If you’ve been with me since Nagios, this is the post where I explain what the pager looks like in 2026.
OpenTelemetry’s GenAI moment
By 2024, OpenTelemetry had won the wire-format war. Vendors that fought it for years started shipping OTel-native ingestion. The collector pipeline became universal. The thing that finally unified metrics, logs, and traces into a single semantic model was, almost by accident, in the right place at the right time when LLMs hit production.
The OpenTelemetry community’s response was the GenAI semantic conventions — a set of standardized attributes for spans representing AI model calls. The shape of a gen_ai.* span:
gen_ai.system = "openai" | "anthropic" | "ollama" | ...
gen_ai.request.model = "gpt-4o" | "claude-opus-4-7" | ...
gen_ai.request.temperature
gen_ai.usage.input_tokens
gen_ai.usage.output_tokens
gen_ai.response.id
gen_ai.response.finish_reasons = ["stop"]
gen_ai.operation.name = "chat" | "completion" | "embedding"
That’s a span representing an LLM call as a first-class operation, with token counts, model identity, and cost-relevant fields. It rides the same OTel collector you set up in Part 4. It lands in Tempo, Datadog, Coralogix, or wherever your traces go. The architecture is unchanged. The semantics are new.
This is why I keep emphasizing that the pipe was already built. Teams that did the OTel work in 2022–2024 were instantly ready when GenAI workloads showed up. Teams that didn’t are now doing both projects at once.
[IMAGE_PROMPT: A pipeline diagram. Application code on the left calls “LLM API” via an OTel-instrumented SDK. The span flows through the OTel Collector and fans out to multiple backends — Tempo, Datadog, Coralogix, Phoenix, Langfuse. Highlight the gen_ai.* attributes attached to the span. Caption: “Same pipe. New semantics.”]
What’s different about AI workloads
The conceptual leap from “monitoring services” to “monitoring AI” is bigger than the leap from “monitoring hosts” to “monitoring services” was. Three properties of AI workloads break the deterministic model in ways the cloud never did:
1. Non-determinism is a feature, not a bug. The same input produces different outputs across runs. There is no fixed “correct answer” you can compare against. This isn’t a flaky test — it’s the design. Replay-based debugging, exact-match assertions, deterministic SLOs all stop applying.
2. A 200 OK can still be a hallucination. The HTTP layer reports success. The model returned text. The text is plausible, well-formatted, syntactically valid — and wrong. Every old layer of monitoring says everything is fine. Only a separate evaluation step can catch this.
3. Cost per call is variable and significant. A traditional API call costs essentially nothing per request — the cost is in the infrastructure. An LLM call costs per token, and a single agentic interaction can easily fan out into dozens of model calls. Suddenly a single user request has a per-unit cost you can read off a bill, and that cost varies based on prompt size, context retrieval, model choice, and how many times the agent decided to “think harder.”
Each of these breaks something. The first breaks “is the output correct?” The second breaks “is the service healthy?” The third breaks “is this profitable?” You can’t fix any of them with the old vocabulary.
The new metric vocabulary
Modern LLM observability has its own dimensions. They cluster into four buckets:
Cost & latency
- Tokens in, tokens out, per call and per session.
- Dollar cost per request (model price × tokens).
- Time-to-first-token (TTFT) — the streaming UX metric.
- P95 generation time, P95 end-to-end including retrieval.
Quality
- Groundedness — does the answer cite the retrieved context, or hallucinate?
- Faithfulness — is what’s cited actually in the source documents?
- Context precision — were the retrieved chunks relevant to the question?
- Answer relevance — does the response actually answer what was asked?
- RAG retrieval recall — did we retrieve the documents that contained the answer?
Safety
- Toxicity scores on output.
- PII leakage detection.
- Prompt injection / jailbreak attempt detection.
- Refusal rate (legitimate vs over-refusal).
Drift
- Moving averages on quality scores over time.
- Distribution shift in input characteristics (are users asking different things?).
- Embedding drift (has the document corpus changed under us?).
The drift bucket is the one that maps cleanly to the old monitoring world. It’s the new “service degraded.” When your groundedness score drops from 0.91 to 0.74 over a week, something has changed — a new model version, a corpus update, a prompt template change, a shift in user behavior. You don’t always know which. You do know the signal is real.
Instrumenting with OpenLIT
Hand-instrumenting LLM calls with the raw OTel SDK is doable but tedious. You’d wrap every model client, capture inputs and outputs, attach gen_ai.* attributes, manage span hierarchy across multi-step agents. For a serious application, that’s hundreds of lines of cross-cutting code.
OpenLIT is one of several projects that solved this with auto-instrumentation. Drop-in OTel-native instrumentation for the LLM layer — OpenAI, Anthropic, Cohere, Mistral, Hugging Face, vector databases (Pinecone, Chroma, Weaviate), frameworks (LangChain, LlamaIndex). One import, full GenAI spans flowing into whatever backend you already have.
import openlit
openlit.init(
otlp_endpoint="http://otel-collector:4318",
application_name="support-agent",
)
# Existing code unchanged. Every LLM call now produces gen_ai spans
# with token counts, costs, latency, and prompts/completions.
import openai
client = openai.OpenAI()
response = client.chat.completions.create(
model="gpt-4o",
messages=[{"role": "user", "content": "..."}],
)
That’s the whole instrumentation story for the basic case. The spans land in Tempo or Datadog, the metrics derived from them land in Prometheus or Mimir. Grafana dashboards work. Alertmanager works. The Part 2 stack is still doing its job — it’s just observing a new kind of workload.
The honest caveat: auto-instrumentation captures the obvious axes (tokens, cost, latency, model, finish reason). The quality axes — groundedness, faithfulness — require evaluation, which is a separate problem. OpenLIT integrates with evaluation frameworks; it doesn’t replace them.
Other projects worth knowing in this space:
- Langfuse — open-source LLM observability with a strong eval focus.
- Phoenix (Arize) — open-source LLM tracing and eval, with notebooks-first DX.
- Helicone — proxy-based LLM observability, drop-in via base URL change.
- Datadog LLM Observability, Coralogix AI — the hosted players extending into this space.
The space is moving fast. The convergence point is OpenTelemetry’s GenAI conventions — pick a tool that speaks them, and you keep optionality.
LLM-as-a-Judge and trace-based evaluation
Quality metrics — the “did the model actually do its job” axes — can’t be computed from the trace alone. The trace tells you what happened; you need an evaluator to tell you whether what happened was good.
The dominant pattern is LLM-as-a-Judge: sample a fraction of production traces, send them to a separate evaluator model with a structured rubric, get back a numeric score, feed that score back into your metrics pipeline as a gauge. Loop closed.
A simplified shape:
def evaluate_groundedness(question, retrieved_context, generated_answer):
rubric = """
Score the answer's groundedness 0.0 to 1.0.
1.0 = every claim in the answer is directly supported by the context.
0.0 = the answer contradicts or invents claims not in the context.
Question: {question}
Context: {context}
Answer: {answer}
Respond with a single float.
"""
score = judge_model.generate(rubric.format(...))
return float(score)
# In a trace-aware sampler:
if random.random() < 0.05: # 5% sampling
score = evaluate_groundedness(question, context, answer)
metrics.gauge("llm.groundedness", score, tags={"model": model, "version": version})
This is the closing of a loop the M/ADLC from Part 2 always wanted but never had a runtime for. You can now define a metric whose computation requires another inference step, and treat its output as a first-class signal in your monitoring stack. The cost-vs-coverage trade-off is real (every evaluation is another LLM call), which is why sampling and offline eval pipelines matter.
The patterns that are emerging in production:
- Online sampling at 1–10% of traffic for cost reasons, with full eval for flagged cases.
- Offline eval suites run against curated datasets nightly — a regression test for model behavior.
- Comparative eval — A/B testing two prompt templates or two model versions on the same inputs and comparing scores statistically (Part 4’s distribution-comparison thinking, applied to AI quality).
- Human-in-the-loop sampling where the LLM-judge flags ambiguous cases for human review.
This is genuinely new territory. The patterns will keep changing for the next few years. The discipline — instrument, sample, score, alert on drift — is the same discipline we’ve had since Part 2.
[IMAGE_PROMPT: Loop diagram. User request → LLM application (with retrieval, generation, OTel spans) → Trace lands in observability backend → Sampler picks N% → Judge LLM scores → Score lands as a metric → Alertmanager fires on drift → Engineer investigates. Highlight that “Judge LLM” is itself an LLM call subject to the same observability.]
Total cost of request — the real new SLO
Latency was the old north star. Cost-to-resolution is the new one.
Here’s the comparison that matters:
| Approach | Latency | Cost per call | Resolution rate | Cost-to-resolution |
|---|---|---|---|---|
| Single-shot, fast model | 2s | $0.005 | 60% | $0.013 (incl. retries + escalations) |
| Multi-step agent, larger model | 30s | $0.40 | 95% | $0.42 |
| Plus support escalation cost on failure | — | — | — | $40+ per failed case |
A 30-second multi-step agent that resolves a ticket for $0.40 beats a 2-second single-shot at $0.005 if the single-shot fails to resolve and creates a $40 support escalation.
This is a new kind of SLO and it requires a new kind of observability:
- Resolution rate — did the user actually accomplish what they came to do? Hard to measure; usually inferred from follow-up behavior, explicit feedback, or downstream signals.
- Cost per resolved request — total token spend divided by successful resolutions.
- Escalation rate — what fraction of cases bounced to a human or to a more expensive path?
- Tool call efficiency — for agentic systems, how many tool calls did the agent make per resolved task? Is that going up or down over time?
I’ve started thinking of this as FinOps for AI — the same discipline that emerged for cloud spend, applied to inference spend. The same dashboards. The same question: “are we spending money on outcomes or on motion?”
The teams that get this right wire cost into their dashboards from day one. The teams that don’t get a finance call in month three.
Closing the arc
Twenty-five years.
Slide 20 of the original “Modern Monitoring” deck told us to let Nagios die peacefully. That was 2014. A decade and change later, the question isn’t “is the host up?” — it’s “is the agent still trustworthy, and can we afford it?”
The pager moved:
- From
exit 2(Part 1) — the script said no - To
error_rate > 1%(Part 2) — the threshold was crossed - To “we can’t query a year of data without an extra zero on the bill” (Part 3) — the cost shape changed
- To “the canary’s distribution is regressing” (Part 4) — the populations differ
- To “groundedness dropped from 0.9 to 0.7, and we’re at $3.20/query” (Part 5) — the answer is unreliable, and expensive
Each shift was a new noun. The host. The service. The fleet. The request. The intent.
The deterministic stack didn’t disappear. Nagios still runs in places where the assumptions still hold — small static fleets, predictable workloads, binary failure modes. The dimensional stack — Prometheus, Grafana, Alertmanager — is still the backbone of every modern platform. The cloud-native scaling stack — Thanos, VictoriaMetrics, Loki, Tempo — is still how you run observability at scale. eBPF and progressive delivery are still how you ship code safely.
What’s new is the layer on top: probabilistic observability for systems whose correctness is itself probabilistic. Confidence scores. Cost per resolution. Drift detection. LLM-as-a-Judge.
If you ran Nagios in 2010 and you’re instrumenting an agentic system in 2026, congratulations — you’ve watched the entire arc. The principles you learned holding a pager at 4 AM are still the principles. Alerts are a contract. Severity matters. Coverage isn’t the goal — actionability is. The vocabulary expanded. The discipline stayed.
That’s the series. Thanks for reading.
[IMAGE_PROMPT: The full evolution table from Parts 1 through 5, rendered as a single landscape graphic. Five eras across the top — Deterministic, Dimensional, Cloud-Native at Scale, Journey-Aware, Probabilistic. Below each: the unit of work, what success looks like, what failure looks like, recovery model. Use a gradient progressing from blue to teal to purple across the columns to suggest evolution. This is the visual that anchors the whole series.]
Resources
- OpenTelemetry GenAI semantic conventions — opentelemetry.io/docs/specs/semconv/gen-ai
- OpenLIT — github.com/openlit/openlit
- Langfuse — langfuse.com
- Arize Phoenix — github.com/Arize-ai/phoenix
- Helicone — helicone.ai
- Ragas (RAG evaluation framework) — github.com/explodinggradients/ragas
- The original “Modern Monitoring” deck — the source for Parts 1 and 2 of this series
The series, in full
- Part 1: Exit Codes and Static Hosts — The World That Built Us
- Part 2: Tags, Time-Series, and the Metric Lifecycle
- Part 3: When Prometheus Outgrows One Box — Cloud-Native Storage and the Scaling Wars
- Part 4: From Metrics to Movement — eBPF, the User Journey, and Progressive Delivery
- Part 5: Confidence Scores and Cost-Per-Request — Observing the AI Stack (here)
Discussion