An exposure sketch, not a design. The plumbing landed first (an all-batches-failed
job is now failed and classified; the refusal mapper is wired back up;
the Mac out-of-credit pill hears about sidecar jobs). What is still open is
where availability should show, and how hard it should push. The argument
below is that the honest answer is three tiers, because the three states are known
at three different moments and with three different confidences — one global
“LLM unavailable” switch would have to pick one and be wrong about the
other two.
The distinction that matters is when the answer is knowable and how long it stays true. Configuration is a local fact and never stale. Credit is a fact about someone else’s server, and is stale the instant it is read. What actually happened is only knowable after spending. None of the three earns a change to the control, and the table’s last column is why: each argument for one turns out to restate a global condition locally, on every card on screen, to save a click that is either free (the pre-flight refuses without spending) or is itself the only accurate test.
| State | Known when | Stays true? | Treatment | Why not the obvious thing |
|---|---|---|---|---|
| No provider no key · local-only · ambiguous |
Before the click, free | Yes — local config | Nothing on the control. The refusal costs no tokens and names the fix in the toast. | Not a replaced control (“Set up AI…”) and not a disabled one. Both restate a global condition on every card that happens to be on screen, to buy back a click that costs nothing — the pre-flight refuses before spending. |
| Out of credit last verdict said dry |
Only by spending (cached ≤24h) |
No — a top-up happens in a browser tab we cannot see | Nothing on the control. The pill carries it, once. The click is the probe; the toast is the report. | Not a block, and not a warning either. OutOfCreditModel
documents that its cache lingers after a top-up, so both would fire at a
researcher who has already paid — one stopping them, one talking them out
of it. |
| Refused now the call actually failed |
After spending | It is the ground truth | Fail loudly, name it, record it. | Not “partial success”. This is the tier that produced the Tagged 0 of 33 screenshot.plumbed |
Not two ways and not three — one. Install stays Install in every state. It is not replaced by “Set up AI…”, it is not disabled, and it grows no warning. The researcher clicks, and the toast says what happened. Everything below is the same card and the same button; only the outcome differs.
The reasoning is one line: this is a global condition, and a codebook card is a local surface. A screen of twenty cards would say the same sentence twenty times, in the place least able to act on it, built on a cached verdict that is wrong the moment the researcher pays. It is already stated once in the titlebar, where a standing condition belongs, and it will be stated again — accurately, and within seconds — by the toast.
The controlOne card. Unchanged in every state below.
Install is apply (D4). It spends, with no confirmation, because the researcher asked by clicking.
Outcome A — it worksThe overwhelmingly common case, and the one the other two must not tax.
Any pre-emptive check would have to run, and be paid for or be stale, on every one of these.
Outcome B — no providerThe server refuses before spending anything; the refusal names itself on the wire.
Verbatim from the en locale, chosen by
autocodeRefusal() from the stable no_api_key reason
— not the wire detail, which told a Mac-app user to edit a
.env file.plumbed
Outcome C — out of creditFound out by trying, which is the only way it can be found out.
And this verdict is what lights the titlebar pill, so the standing condition gets stated once, in one place, from the only event that actually established it.plumbed
Outcomes B and C both end with the codebook installed and uncoded:
importCodebookTemplate runs before startAutoCode, so the
groups are linked whatever the job then does. That state is not new — it is the
consequence of never gating, accepted deliberately — but it is currently
invisible and unrecoverable, and that is the real work. Section 5.
Before and after the plumbing. The left chip is the 3 Sep screenshot; the right is what the same failure produces now. Nothing here needs designing — it is shown because it is the load-bearing surface: once the control never changes, the toast is the only thing that tells a researcher what happened, so it has to be right.
BeforeAll 33 quotes refused, reported as a partial success, offering a report that was empty — and the link doubled as the dismissal, so the chip could not be closed.
Cause: gather(return_exceptions=True) swallowed every
batch error, so the classifier — which knows out-of-credit from rate-limited —
was never reached.
AfterNamed, dismissable, and no door to an empty room.
The sentence comes from autocodeFailure(), which
already had it. The only thing that changed is that failure_kind
now reaches it.plumbed
The pill is settled and is not redrawn here — it shipped from
out-of-credit-ux.html,
which is the design of record. Reproduced below only so tier 2 has a referent.
What changed is upstream of it: OutOfCreditModel was fed from
PipelineRunner.deriveFailureState alone, so a sidecar AutoCode job
that emptied the account lit nothing. It now hears via the bridge.
The pill is the whole of tier 2’s ambient presence — one place, and
its popover is the only resolve surface.
Shipped, unchangedToolbar .status zone,
alongside the Ollama pill in the same StatusPill envelope. Amber
dot, never red — red is reserved for a genuinely failed run.
Nothing goes in the content area. Add funds… opens
the provider’s billing console; Switch provider… deep-links
Settings ▸ AI. Secondary leads, default trails — macOS puts the
default rightmost, which is why the button order here is the reverse of the
original mockup’s.
Clears when a fresh validation records .online. Known limitation,
unchanged: refocus re-reads the cache without revalidating, so it lingers
after a browser top-up until Settings next checks.
Every server route that spends tokens, and what it does today. With the control never changing, each of these needs exactly two things and no more: a pre-flight that refuses before spending where it cheaply can, and a failure that arrives named. No per-surface availability state, nothing to keep fresh. Three of the five have neither.
| Surface | Route | Tier 1 gate | Tier 3 message | Note |
|---|---|---|---|---|
| Install / AutoCode | routes/autocode.py |
yes — 3 refusals | namedfixed | Refusal mapper had no caller since the v1 lens was deleted; now wired. |
| Chat lens | routes/chat_lens.py |
none | raw | 502 carries f"{type(exc).__name__}: {exc}" — the raw SDK text. |
| Codebook builder | routes/codebook_builder.py |
none | unhandled | Calls the LLM directly from the route. |
| Signal elaboration | server/elaboration.py |
silent skip | silent | Logs a warning, returns cached only. The researcher sees fewer elaborations and is told nothing. |
| Analyse / re-analyse | routes/analysis.py |
none in serve | events | The CLI path has preflight_api_key; serve does not. |
Two hand-rolled copies of _has_api_key exist
(routes/autocode.py, elaboration.py), neither reading the
provider registry — a third provider-name switch to forget when a provider is
added. One llm_availability(settings) → (ok, RefusalReason)
collapses both and gives the four ungated surfaces the same vocabulary AutoCode
already speaks.not done
Never gating means outcomes B and C leave a codebook installed and uncoded. That is accepted. What is not yet decided is how the researcher gets out of it once they fix the provider — by switching to one that works, or by topping up the one that is stuck. The expectation to design against: autotagging should come back to life on its own.
| Piece | State | Detail |
|---|---|---|
| Retry is possible | works | start_autocode_job 409s already_applied on a
completed job but discards a failed one and re-runs. Calling the
all-failed case completed was what walled the framework off for
good.fixed |
| A retry codes everything | works | completed_at is the applied watermark, not a stop time. A job
that coded nothing no longer stamps one, so it stays out of
reapply_active_frameworks’ maintained set and a retry sees
the whole corpus rather than the delta since the
failure.fixed |
| The state is visible | no | installed means “has groups in this project”, and
the template is imported before the job runs — so the card reads
Uninstall and the navigator draws an uncoded codebook exactly like a
coded one. Recovery today is Uninstall → Install, undiscoverable. |
| Something triggers the retry | no | Nothing re-runs a failed job. This is the design question. |
| Event | What we get | Use |
|---|---|---|
| Switch provider, or paste a new key | .bristlenosePrefsChanged →
ServeManager.restartIfRunning() |
Strong — the sidecar restarts, so there is a clean moment to act. |
A validation records .online after .outOfCredit |
verdict-cache transition | The truest trigger — it fires on the change itself. But only when something actually revalidates. |
| Adding credit | nothing, ever | It happens in a browser tab. OutOfCreditModel’s own
comment records that refocus re-reads the cache without revalidating —
which is also why the pill lingers after a top-up. |
.outOfCredit → .online, re-run whatever failed —
sequentially, as reapply_active_frameworks already does “so
spend stays legible”, with the ordinary progress chips. This is the
jump-back-into-life, driven by the real event rather than a timer, and it fixes
the pill’s lingering as a side effect.preflight/api_key.py already
spends ~$0.0001 to answer this for run/analyze. Its
only use here was the pre-emptive warning, which is rejected — so it would charge
the researcher for a question nobody asks, on a surface they may only be reading,
and still not answer for two providers. The run is the probe.