TL;DR
I added a Jev-style decision gate to an Argo Rollouts canary: one yes/no question, one probability, and thresholds that map to Promote, Abort, or Argo’s built-in Inconclusive pause. Then I held each canary and asked the same question ten times. The healthy and broken releases were boring, as they should be. The interesting one was a release that was 1 second slower with zero errors. The classic success-rate check passed it. The model said “block” on 7 of 8 clean calls. On the one call Argo actually used, it said 0.1, and the release shipped. A gate isn’t one answer from a smart model; it’s a distribution, and you have to design for that.
Originally published at portfolio.hagzag.com.
13:20:36, Blue Goes to 100%
This is the line that made me rewrite this post:
13:20:32 canary success=1.000 p95=1005ms stable_p95=4ms p_block=0.9 decision_ms=1167 # my 10th probe
13:20:35 canary success=1.000 p95=1007ms stable_p95=6ms p_block=skipped # Argo: success-rate metric
13:20:36 canary success=1.000 p95=1004ms stable_p95=5ms p_block=0.1 decision_ms=1036 # Argo: jev-decision metric
Same canary, same model, same question, four seconds apart. Seconds earlier the model had said 0.9 (“block this release”) seven times out of eight. Then Argo asked once, got 0.1, and promoted a release that was 200 times slower than the one it replaced.
In my last post I argued that most of an agent’s “thinking” is really a multiple-choice question, and that a confident witness is not a judge. This time I wanted to see what that looks like inside a real release pipeline, so I built one.
The Gate
The lab is public: hagzag/jev-rollouts-demo. It’s a k3d cluster with Argo Rollouts and the official, unmodified argoproj/rollouts-demo image. The image already has the three behaviors I needed, selected by tag or env var:
| Release | How | Behavior |
|---|---|---|
| green | argoproj/rollouts-demo:green | healthy |
| red | argoproj/rollouts-demo:bad-red | ~85% HTTP 500 (ERROR_RATE=15, which is actually a success rate) |
| blue | argoproj/rollouts-demo:blue + LATENCY=1 | one extra second per request, zero errors |
The canary goes to 20%, pauses, then runs an AnalysisTemplate with two web metrics against a small Go service, jev-gate. The gate probes the stable and canary Services, turns what it saw into JSON state (success rate, p50/p95, a few sample responses, the SLO), and asks a decision service one noul question in the OpenJev/Jev wire format:
The canary release should be blocked from a full rollout because it would noticeably degrade the user experience compared with the stable release.
The decision service was the hosted OpenJev with gpt-4o-mini behind it. That’s an open approximation of TypeSafe’s Jev, not the real model. The demo data has no PII, so a hosted service was fine here.
The thresholds are where Argo does the interesting work:
- name: jev-decision
count: 1
successCondition: result < 0.3 # promote
failureCondition: result > 0.7 # abort and roll back
# 0.3–0.7: neither matches -> Argo marks the measurement Inconclusive
# and pauses the rollout for a human
provider:
web:
url: "{{args.gate-url}}?canary=...&stable=..."
jsonPath: "{$.p_block}"
That middle band was the whole idea: a calibrated “I’m not sure” becomes Paused - InconclusiveAnalysisRun, a state Argo already supports.
Ten Questions Per Canary
For the evidence run, a script resets the stable version, releases each color, waits until the canary is serving, holds the rollout, asks the gate ten times, then lets Argo decide on its own. Results (calls where the canary Service still pointed at stable, fully or partly, are excluded; more on that below):
| Release | Classic gate (success rate) | p_block over clean calls | Would have done | Argo’s own call | Outcome |
|---|---|---|---|---|---|
| green | 1.0 | 0 ×8, 0.01, 0.1 | Promote 10/10 | 0.1 | Healthy |
| red | 0.05–0.3 | 0.9 ×7, 0.95 | Abort 8/8 | 0.9 (success 0.05) | Degraded (rolled back) |
| blue | 1.0 | 0.9 ×7, 0.1 ×1 | Abort 7/8, Promote 1/8 | 0.1 | Healthy (shipped) |
Green and red are boring in the right way. Red didn’t need a model at all: a 5% success rate fails the classic metric on its own, and the model just agreed.
Blue is the case the gate exists for. The success-rate metric reported Successful with value=1, because a request that takes a second still returns 200. The model saw the problem most of the time. It just didn’t see it the one time it counted.
What the Numbers Say
The answers come in buckets. Across 30 calls the model returned exactly six distinct values: 0, 0.01, 0.1, 0.8, 0.9 and 0.95. Nothing between 0.1 and 0.8. A chat model asked for “a probability” picks a round number that sounds right, so the Inconclusive band I designed around was never used once. The Argo feature worked; the model never gave it a reason to fire.
Uncertainty showed up as flips, not as middle values. On blue, the honest answer is “this is debatable”: inside the 1.5s SLO I gave it, but 200× slower than stable. A calibrated model would say something like 0.5. This one said 0.9, 0.9, 0.9, 0.1, 0.9… The uncertainty is real; it just leaks out as variance between calls instead of a number you can threshold.
It isn’t fast. Decisions took 818–2,727 ms (median 1,053 ms) through the hosted wrapper. Fine for a canary analysis that runs once per release. Nowhere near the 70–500 ms TypeSafe quotes for real Jev, and not something I’d put on a request path.
Eight samples is not a study. These are 8 clean blue calls on one afternoon. I’m not claiming gpt-4o-mini says “block” 87.5% of the time. I’m claiming that a single call is a sample, and treating a sample as a verdict is how blue shipped.

What Actually Went Wrong
My first evidence run measured stable twice. I waited until the new pod showed up in the canary Service’s endpoints and started probing. But Argo hadn’t switched the Service selector yet, so the first one or two calls in red and blue hit the old version: success=1.000 p95=2ms for a release that was mostly 500s. Those calls are excluded above, and the script now waits for the selector itself plus a few seconds. Without a reset step, my very first “green” run was also a no-op: stable was already green, so nothing rolled out and the script happily reported an old AnalysisRun.
The same canary had already flipped once. Before the scripted run, a manual blue release with the identical pod template got p_block=0.9 and was aborted. Eight minutes later the scripted one got 0.1 and promoted. If I’d only run it once, I’d have written either “the model catches slow releases” or “the model misses slow releases”, and both would have been wrong.
What I’d Change Before Trusting It
Argo already has the knobs; I just started with count: 1. Next version of the metric:
- name: jev-decision
count: 5 # ask five times, 5s apart
interval: 5s
failureLimit: 2 # 3+ "block" answers -> Failed -> abort
inconclusiveLimit: 2
successCondition: result < 0.3
failureCondition: result > 0.7

That turns one sample into a vote, using nothing but Argo. Beyond that:
- Give the model the comparison, not just the raw numbers. I passed the SLO and both p95s; I never passed “canary is 200× slower than stable” as a field. The model anchored on the SLO some of the time.
- Keep the deterministic gates. Success rate caught red faster and for free. The model’s job is the gray zone, not everything.
- Measure calibration on your own history before you trust a threshold. Replay the last 50 releases through the gate, compare
p_blockwith what humans decided, and only then pick 0.3 and 0.7. A trained decision model (TypeSafe’s Jev, or open ones like RSI-Jev) is supposed to fix the bucket problem. “Supposed to” is exactly what this lab is for.
Conclusion
The demo did what I built it to do: three releases, three different outcomes, and a gate that saw a regression a success-rate metric can’t see. It also did something I didn’t plan for: it showed me that one model call is a coin with a weighted edge, and my release pipeline flipped it once.
The fix isn’t a smarter model. It’s treating every model answer as a measurement, and asking more than once before you bet production on it.
Discussion