Index of designs
Every panel on this page carries a code, so any of them can be named without
being described. V is a row layout, T a
treatment of the Codebooks section, J and D a
whole navigation; P codes are data, not design. A project panel
is stamped T4·P1 — treatment and project — so a
screenshot of one says what it is on its own. Every code below is a
link: it scrolls to the panel it names and flashes it. A
T code switches the Codebooks-section control first, because a
treatment is a mode rather than a panel sitting there waiting to be scrolled to.
Decision trail
What was chosen, what was abandoned, and the evidence that moved each one. Panels on this page carry their status in the title bar — chosen panels take a green edge, abandoned ones fade, current means what the product does today, which is not the same as chosen. Every code below is a link to the panel it names.
Gaps in the spec
What a build would have had to guess at. Scanned 13 Sep 2026 against the shipped code rather than against this page, which is where it earned its keep — several were places where the mockup and the product quietly disagreed and nobody had chosen. All of them are now closed: six became decisions in the trail above, and the two below turned out not to be questions.
- Where do the two heatmaps go? — nowhere, they never moved. They are
sources for
InspectorPanel, a sibling of.analysis-center, so they are not in the cards column at all and the refactor does not touch them. Cells go on callinghandleCellClickand scrolling to cards. Their own defects are real and deferred, below. - Empty state — not a design question.
AnalysisPagealready has a!hasSentiment && !hasTagsbranch renderinganalysis.noData, andAnalysisSidebarreturningnullunder the same condition is right: the page is already saying there is nothing, so the sidebar should not say it twice. And zero tags from N quotes from real transcripts means something upstream failed — that is a pipeline problem wearing a lens costume, not a state to design for. There is no partial case, because locations are derived from cards.
Deferred by decision
Not gaps — choices to stop here, with somewhere to come back to.
- Where signal cards are built, and what the assistant sees. The lens
builds them in the browser \u2014 two endpoints merged in
AnalysisPage, with de-duplication at the merge point, because that is the only place both kinds exist together. Tactical and it ships. But there is a second answer to what the signal cards are:load_signalsinserver/grounding.pyserves both the MCPget_signalstool and the chat lens, and by design it computes over the curated corpus — hidden quotes out, edits substituted, and unreviewed AutoCode proposals contributing nothing. The principle is right. The size of the gap had never been measured: on P1 the lens draws 29 cards across eight codebook groups andload_signalsreturns 6, every one of them Sentiment, because 112 of 143 tag assignments are proposals nobody has reviewed. An agent asking what the signals are is told about sentiment and nothing else, confidently, while the researcher looks at eight groups. Deferred with the server-side question, on the board. - The view menu. Hidden today \u2014 an idea not fully realised. Strongest signal and by-codebook stay drawn as U1 and U2 and stay in the code; sections and themes is the only order that ships. Shipping the menu on day one is how “one better lens” quietly becomes a menu of equals, and the existing order is meant to be insurance rather than a co-equal choice.
- Normalising signal strength. Handed to a session of its own,
docs/design-signal-strength.md. Until it lands, ordering is directionally right and the number is not comparable, and the Sentiment card keeps its exemption from de-duplication. - A signal’s valence.
success/gap/tension/recoverykeep being computed and stored; none of it reaches the screen. The all-caps chip was the failure, not the classification. A second, undisplayed vocabulary (classify_flag’s Win/Problem/Niggle/Success/Surprising) should be settled in the same pass. - Elaboration coverage. 18 of 76 surviving cards have no name, and 14 of 60 locations have none at all. Shipping with bare chips by decision. The fix is cheap after de-duplication — 18 elaborations across seven projects — but it needs a prompt for thin evidence, not a bigger number.
- Floors and thresholds.
MIN_QUOTES_PER_CELL = 2is a volume floor and two quotes is barely a signal from a five-participant study. Settled by checking real interview data against outputs, not from this page. - Masonry for the signal cards — PARKED 20 Sep 2026, and the plan
with it. The Quotes lens tessellates its cards with native CSS masonry:
display: grid-lanesbehind@supports, no JS and no fallback code, engines without it keeping the rectangular grid (organisms/responsive-grid.css;.codebook-gridis the second adopter, and puts its section headers inside the grid atgrid-column: 1 / -1so lanes reset per section). A plan to bring it to.signal-cardswas written on 14 Sep and is superseded rather than deferred — do not pick it up as written. It assumed a multi-column grid:repeat(auto-fill, minmax(min(100%, 26rem), 1fr))with--bn-grid-gap, measured at 1–4 columns and card heights from 230 to 1442px. That is gone..signal-cardsis nowgrid-template-columns: 1fr; gap: 0— a single joined run whose card corners come from position in the run, because a location with one card is the common case. Masonry against one column with no gap has nothing to pack. It was already the hard one under the old shape, which is worth carrying forward: the location grouping fragments the list into a grid per place, 85% of them holding a single card, so lanes had almost nothing to work with; and Chromium reportsCSS.supports('display','grid-lanes')false, so Playwright can only ever verify the fallback and the masonry itself is checkable in WebKit alone. If the tessellated look is ever wanted here, start from the run-with-seams the cards have now and ask whether it should be a grid at all — not from this plan. - The heatmaps. Good enough for now, improved in a later release by
decision. Known and accepted: a cell can be drawn
.has-card, accept a click and silently do nothing, becausesignalKeysis built from the uncapped signal lists whilecardRefsonly ever holds rendered cards — so every cell outside the top six is already a dead click. De-duplication adds a second class of the same thing. Whatever hides a card should eventually un-hot its cell. - Quote-level deep links. A card’s location now links to that place’s anchor in the Quotes lens, on both card kinds. Linking an individual quote is still open.
Latest iterations
The designs still under consideration, lifted to the top of the page: N, the latest drawing, and the three benches that carry a chosen panel — U the navigation, C the two card kinds, S what a card is. Everything below them is the discussion that produced them, unchanged and in the order it happened.
N · the latest thinking, drawn
Everything the trail marks chosen, applied at once: the navigation by location with T6b rows, and both card kinds in both states. The two cards are the same design — the only difference is whether the chip names a sentiment or a tag group, which is what “a card is a location × a tag group” buys.
Real throughout: locations, group names, tags present, timecodes, participant ids,
intensities, and every metric, read at MIN_QUOTES_PER_CELL = 2.
Not real, and said here rather than buried: quote text is
synthetic at each real quote’s measured length, and the two elaborations are
rewritten into the two-sentence grammar — the stored ones are still single
em-dash-joined sentences, so this is what the prompt change would produce, not what
it has produced.
Look at N3 beside N5. The sentiment card’s Conc. reads 1.0× and the codebook card’s 2.2×, and the first is a structural constant — a one-group framework has a one-column matrix, so concentration cannot say anything there. The two numbers in the hero chips are not yet measured the same way, which is the open normalisation question showing up in the drawing.
U · a view menu, and one hit area per signal card
Under Beds & Mattresses Category alone, P1 holds eleven distinct signal cards — two from the sentiment matrix and nine from the codebook matrix, from four different frameworks. Every one of them is a place the sidebar should be able to send you. T5, T6 and T6b do emit a link per signal, but the enclosing place inherits the row hover, so the whole block lights up and the individual targets are invisible. Here nothing but a signal is a link — a place or a codebook is a heading, not a control — and the commentary asserts it: links total should equal signal cards.
The menu at the top of each panel is the organising principle itself, which is the thing being tried. Sentiment appears in U2 under its own codebook heading, because it is one. All three follow the strength floor and signal budget.
C · normalising the two card kinds
The main content area draws two different cards for what is structurally the same thing. A codebook signal has an elaborated name, so the name is the heading, the location drops to the eyebrow and the badges move to the top right. A sentiment signal has no name, so the location is the heading and the badge sits under it. Side by side they share almost no structure.
The summary a sentiment card needs already exists. “Category Navigation Confusion”, and a written elaboration to go with it, is sitting on the Sentiment-group codebook row for that same location — generated because sentiment is a codebook and gets elaborated like one. The sentiment matrix cell, which is what the card actually renders, has none. So the two halves of this session’s argument turn out to be a single move: stop drawing the Sentiment-group row as its own card, and fold its name onto the sentiment card that has none.
Metrics, timecodes, participant ids, intensities and tags are the real P1 values. Quote text is synthetic, written to the measured character length of each real quote — this file is tracked and participant speech is not, and length is what the layout is judged on.
S · one rule for what a card is
A card is a location × a tag group, and it surfaces that group’s
tags. Structure carries interaction design,
information architecture, navigation pattern. Under that
rule sentiment is a group like any other, and confusion and
frustration are tags inside the Sentiment card — not cards of
their own.
Today the lens breaks the rule in exactly one place: sentiment is also analysed at tag level, so every sentiment value competes for a card. That is the whole asymmetry between the two card kinds, and on this project it earns nothing. Where a sentiment value clears the two-quote floor, the place states its sentiment twice — once as a bare value with no summary, once inside the Sentiment card. Where no single value clears it, the tag-level analysis contributes no card at all: on three of these five places, including Shopping Bag and Search & Sort Results, every sentiment value falls below the floor on its own while the group clears it comfortably on three and four quotes.
So across five places the tag-level analysis adds exactly one card, and that card is a duplicate of something the Sentiment card already says — and says better, because the group card has a written summary and the bare value does not.
Read at MIN_QUOTES_PER_CELL = 2, the shipped floor, so
every card here is one the product would really draw. Nothing on this bench is
invented.
R · one card per set of quotes
Once a card is a location × tag group, a place with five codebooks
installed draws the same quotes several times over. Under
Beds & Mattresses Category the app computes
nine cards from seven quotes by two participants; six of the nine
sit exactly on the two-quote floor, and two of them
— Status visibility and Feedback — are the
same two quotes, named once in Nielsen’s vocabulary and once in
Norman’s.
The rule. Walk a location’s cards strongest first and keep a card only if it brings a quote none of the kept cards already carries. Ties break on quote count, so the card covering more wins. Attention is the expensive thing: asking a researcher to read the same two quotes five times under five headings spends it for nothing, and the answer they reach the fifth time is the one they reached the first. You would not duplicate a sticky across five clusters on a whiteboard — you put it in the strongest group under the strongest interpretation.
Measured across nine real projects — 74 locations, 106 cards: the rule hides 21, and costs no evidence at all, because every quote keeps a card pointing at it. It is entirely a multi-codebook effect: every single-codebook project in the corpus already draws exactly one card per location and loses nothing.
R3 is why the Sentiment card is exempt. Sentiment is a one-group framework, so its concentration is structurally 1.00 and its composite is handicapped against every codebook card. Run the rule unguarded and the Sentiment card is deleted in 5 of the 74 locations — in Shopping Bag it is the second-strongest card present. The exemption is a guard, not a fix; the fix is normalisation, which is open.
W · the main content, in the navigation’s order
The lens draws two flat grids today — #signal-cards-sentiment
then #signal-cards-tags — split by kind, each under its
own SectionHeading. The navigation is about to be organised by
place, so the sidebar would stop being an index of the page beneath it:
same cards, different shape.
W2 is one run of locations in the navigation’s own order,
each under a heading, cards ranked within. Nothing else moves. The heading is
.analysis-codebook-heading, which already exists and already carries
exactly this weight of statement; no new heading style is invented. The card is
untouched — it is drawn compact here because its interior was settled in
section N and repeating it would only make this panel taller.
The heatmaps do not move: they are sources for
InspectorPanel, a sibling of .analysis-center, so they
are not in this column at all and their cells go on scrolling to cards.
Real, from the ikea project at MIN_QUOTES_PER_CELL = 2, after
de-duplication: 19 cards over 10 locations, against the 12 the
shipped MAX_SIGNALS = 6 allows. Sections and themes interleave
unlabelled — the themes are at positions 4, 5 and 8.
Archive of discussion
The working, in its original sequence: the ten row layouts against the width slider, V9 on six real projects, the treatment comparison on one project, the whole-navigation and depth benches, and the measurement essays. Nothing here is discarded — a current panel is what the product draws today, and the abandoned ones are kept so that a settled question is not reopened by accident.
V9 on six real projects
One variant, six sidebars. Each panel is what that project’s Analysis lens
actually renders — read from its own database through the app’s code path
(get_sentiment_analysis and
get_codebook_analysis), sorted by composite_signal
and cut to MAX_SIGNALS = 6 per section exactly as
AnalysisPage.tsx does. Nothing is picked for effect: these are
the strongest six of each, in order, with their real labels, their real
elaborated names, and the codes actually present on their quotes.
Panel width and tags per row drive these panels too; the content buttons do not — each panel carries its own project. Codebooks section switches the treatment:
- as shipped — both sections exactly as the lens renders them today.
- hide Sentiment — cut to six first, then drop the Sentiment-group rows. Nothing moves up, so the gap is what is actually lost.
- hide + backfill — drop them first, then cut to six, so the freed slots refill from the same ranking the analysis already produced.
- one flat list — no headings at all. Every signal the project produced, sentiment and codebook interleaved, strongest first, cut where the strength floor says the list stops being worth reading.
- by location — one row per place, every signal about it a chip you can click. Sentiment chips keep their hue; a codebook signal is named by its group, tinted with that group's colour.
- by location · named — same grouping, but each codebook signal gets its own sub-row carrying its elaborated name, its pattern, and its codes.
Codebook group colours ux emo task trust opp — each group cycles its own slots, so two codes in one group are two tints of one hue. Hover a pill for its group.
The real test: one project, every treatment
Five of the six projects above cannot tell these treatments apart, because they have no codebook tags that are not sentiments — every treatment renders them identically. project-ikea is the only real test. Four codebooks, fourteen groups, 23 non-Sentiment signals, and one section carrying ten signals at once. Same width, cap and floor controls; the Codebooks-section buttons do not apply here — each panel pins its own.
Three jobs, three navigations
project-ikea, one fixed analysis — the same signals, order and strengths in every panel. Only the structure changes. Each of these is drawn for a different answer to “what is the left-hand panel for”, because the three answers do not want the same thing: an index wants targets, a summary wants sentences, and triage wants a comparison. Scrolling is free here, so nothing is shortened to fit.
How far down the nav carries you
The same question held at one job — index — so the depth axis can be read on its own. The tension is the busiest place: Beds & Mattresses Category carries ten signals, and what a structure does with that one entry is most of what distinguishes these. D1 is J1 — the same renderer, repeated so the axis reads without scrolling back.
Both nav benches are P1 at the ≥1 sentiment volume, baked in. They do not follow the sentiment-volume or Codebooks-section controls; the panel width, cap and appearance controls still apply.
Width cannot fix the single-line row
Each V0 panel prints how many characters fit on one line at the current width, and what share of the 347 real labels in the local corpus that covers — 27 project databases, section names, theme labels and elaborated signal names together. Drag the slider and watch it:
| width | fits | corpus covered |
|---|---|---|
| 200px min | 12 chars | 23% |
| 240px now | 18 chars | 51% |
| 280px | 24 chars | 75% |
| 320px | 30 chars | 86% |
| 360px | 36 chars | 93% |
| 480px max | 54 chars | 100% |
Covering the corpus on one line costs the entire drag range. Not the default — the maximum, 480px, the widest a user can make the panel at all, on every lens, permanently. 360px still drops 7%. That is the whole argument for spending depth instead: there is no width worth paying that is also a width that works.
The tail has a specific cause, and it is not the LLM signal names. Those are
capped — signal-elaboration.md says “Signal name MUST be 2–4 words”
and the schema repeats it, and the corpus obeys: median 23, max 31. Theme labels
have no number, only the adjective “concise” in
thematic-grouping.md and ThemeGroupItem — and they run
median 21, p90 41, max 51. The field with a number keeps to it;
the field with an adjective does not. Bounding it is a one-line prompt change,
but it would not settle the layout: theme labels are user-editable, so the row
has to survive whatever someone types regardless.
The list is bounded and ranked — which makes depth cheap
Two facts about the shipped list change the arithmetic of everything above, and neither is visible from the sidebar itself.
It is capped at six per section.
MAX_SIGNALS = 6 in AnalysisPage.tsx, applied as
slice(0, MAX_SIGNALS) to the sentiment and tag lists — and the
sidebar reads those same capped arrays, because setAnalysisSignals
is handed the capped ones. So the panel is never more than twelve
rows. Depth is not an unbounded scroll here; it is a bounded, computable
worst case, which is a much weaker argument against wrapping than it first looks.
It is sorted strongest first. Both analysis modules sort
composite_signal descending, and the client sorts again. Signal
weakens top to bottom. The card shows its rank explicitly
(.signal-rank, “#2”); the sidebar shows no number, so
position is the only rank cue it has.
That is a real strike against the regrouping variants, and it lands hardest on the ones I liked most. V4 buckets rows by badge value and V6/V7 pair them by location — both reorder a list whose order is its ranking. Pairing puts a location at the rank of its strongest signal and silently drags the weaker one up with it; grouping abandons rank entirely. On the ikea sentiment set the four rows happen to pair adjacently so nothing moves, which is exactly the kind of luck that hides the problem. V9 inherits it, since it pairs the sentiment half. Any of these needs an answer to “what happened to the ranking?” before it ships — showing the rank number, or accepting that the sidebar is a directory rather than a ranked list.
And there are two caps, both of them the same kind. The
meaningful one — “don't surface a card if the signal is too weak” — does not
exist. MIN_QUOTES_PER_CELL = 2 is a volume floor, not a strength
one, and confidence (strong / moderate / emerging) computes exactly
that judgement, is passed to the client, and gates nothing. Measured across both
axes of every real project, ranked by composite_signal, of the cards
the sidebar surfaces: 2% strong, 5% moderate, 93% emerging —
“emerging” meaning it failed both bars. In ten of eleven projects the navigation
cap never binds at all, because fewer than six signals exist to cut.
And the binding constraint is participants, not quotes. Taking
the one analysed project with real interview-length transcripts (20 sessions,
9 participants, 242 quotes): 167 carry a sentiment, landing in 58 populated
cells at a median of 2 quotes each — 38 of the 58 clear
MIN_QUOTES_PER_CELL, 15 reach 4, and the median cell carries quotes
from exactly 1 participant. Only 10 clear a ≥3-participant bar and
exactly 1 clears “strong”'s ≥5. (An earlier version of this paragraph
said 484 quotes, 334 sentimented, median 4. That database holds the project
twice — project_id 1 and 2, 242 identical quotes
each — so every per-cell count was doubled; “38 cells clear ≥4”
was really 38 cells clearing ≥2, which re-measuring reproduces exactly. The
argument survives and gets stronger, because the real numbers are half as
generous.) Quotes are not scarce; breadth of agreement is —
which is what simpsons_neff exists to measure, and it is the right
thing to be scarce.
That reprices the whole exercise. Scaled to a five-interview study, the same funnel yields something like two to four signals worth surfacing — not twenty, not even six. So a panel built to rank twelve is, on realistic data, showing three rows. Which makes depth very cheap and truncation very expensive: each row is a large share of everything the researcher gets. (Caveats: that project is oral history rather than usability, its quotes are overwhelmingly themed rather than sectioned, and one project is not a corpus. Also: an earlier version of this paragraph counted the section axis only and read 96% / 4% / 0%. The method was wrong; the conclusion was not.)
The shipped layout on a real project
rockclimbing · shipped is not a stress test. It is a screenshot: the twelve signals that project actually renders, 8 participants, 6 per section, with its real theme labels and its real elaborated signal names. At the 240px the panel now opens to, V0 ellipses 11 of those 12 rows. At 320px — well up the drag range — it still ellipses 11. V6 and V8 do the same, because they keep the badge on the name's line.
rockclimbing · long tail is the same project's longest real theme
labels, the ones ranking happened not to surface: “Inclusion, Representation, and
Gym Infrastructure” (49), “Training Structures and Progression Strategies” (46),
“Footwork, Technique, and Movement Fundamentals” (46). Nothing invented, nothing
padded. V0 ellipses 12 of 12. Its codebook rows carry that project's real groups
and codes — Control and freedom with undo available,
emergency exit, destructive action — so V9 is showing
real open-vocabulary tags rather than sentiment words wearing a different hat.
These labels also carry punctuation the synthetic sets did not: an en dash in “The Physical–Mental Duality of Climbing”, a slash in “Unrelated / Out-of-Scope Content”, hyphenated compounds in “Long-Term” and “Out-of-Scope”. Each is a different break opportunity, and they are the reason a wrapping variant behaves better on this set than the character counts alone would predict.
Stress testing that is calibrated, not invented
Two sets push both axes at once, at measured percentiles rather than at made-up extremes:
- stress · p90 — 41-character row labels (theme-label p90) with 4 tags per row (the median group size) at ~23 characters (tag p90).
- stress · max — 51-character labels (theme-label max) with up to 11 tags per row (the largest real group) at up to 32 characters.
The tag names are real, lifted from the corpus:
subcommand-hierarchy-confusion, multi-terminal-context-switching,
Jakob: convention expectation, destructive-action-protection.
Measured over 620 tag names: median 14 characters, p90 23, max 32, and up to 11
per group. That is roughly double a sentiment word, in an open vocabulary, where
sentiment is a closed set of seven running 5–12.
At stress · max, 240px: V0 and V8 clip 6 rows, V6 clips 5, and every variant that gives the name its own line clips nothing. V9 clips nothing and carries 9 distinct badges against V0's 3 — and note these sets deliberately model a project with three codebook groups, the case where the group name is at its most informative. It still loses three to one.
The cap is the whole design decision in V9
Uncapped, the largest real group puts eleven pills under one name. Measured at stress · max, 240px:
2 588px 6 of 28 codes shown
3 655px 9 of 28
5 770px 15 of 28
all 1036px 28 of 28
Uncapped nearly doubles the panel. Three per row costs 67px over two and shows half again as many codes, which is where I would start — but this is a call to make by looking, which is what the control is for.
The overflow indicator is not a new control. The string is the
shipped analysis.more (“+{{count}} more”, already translated in all
21 locales) and the treatment is .cell-tooltip-footer from
analysis.css, where it already means “there are N more than shown”
— in this same lens, on the heatmap cell tooltip. The only invention here is
placement: the shipped use is a footer line beneath a list, and a footer
line per row is a line per row in a sidebar, so it rides inline at the end of the
pill flow instead. If V9 ships, that treatment should be promoted to a shared
atom rather than copied a third time — .cell-tooltip-footer is named
for its only use site. The 2 / 3 / 5 / all buttons are a playground instrument,
not a proposed affordance; the product would carry one cap, not a picker.
The pill is the shipped quote-card tag pill, and that is an argument
rather than a style choice. The sentiment badges above are a
partition: Quote.sentiment is “a single dominant sentiment,
or None”, so those badges genuinely decompose the location and their counts sum.
Codebook tags are a relation — quote_tags is many-to-many,
and analysis.py's own trade-off note says quotes tagged from several
groups “count in each group column”. So the codebook badges are a cover, not a
decomposition. Borrowing the quote card's pill — where badges have always meant
“codes present on this thing” — lets the two sections claim different things
without inventing an element to say so.
V8 is a different axis from the rest
V0–V7 argue about where the badge sits. V8 asks what it says, and keeps V0's geometry exactly. The commentary strip counts distinct badge values in the Codebooks section — the measure of whether that column carries information or merely repeats the heading above it:
V0 … V7 1 distinct badge
V8 3 distinct badges
One is not a column, it is a caption printed once per row. The codebook
badge shows group_name — and measured across every real project
in the corpus, 11 of 11 have exactly one substantive codebook group
once Uncategorised is excluded. So that badge prints the same word
on every row of every real project. pattern does not: 30 tension,
12 gap, 12 success, 8 recovery across 62 elaborated signals, and it varies
inside a single project (project-ikea alone runs 5 tension, 1 gap,
1 success, 1 recovery).
V8 is strictly better than V0 on identical geometry — same height, one fewer truncation on the elaborated set, because “GAP” and “SUCCESS” are shorter than the group name they replace. And on the un-elaborated set it is shorter than V0 (−12px) with zero badges: no pattern exists yet, so nothing is drawn and the name takes the full width exactly when the name is a short location. Absence carries the information.
So V8 composes with the others rather than competing. The honest recommendation is V8's badge inside V1's wrapping — the badge stops lying, and the name stops being clipped. Nothing here measures that pair yet; it is the next panel to draw.
What the first run measured
Numbers below are from ikea · elaborated at 240px — the real content that ships once elaboration has run, at the width the panel now opens to. They are reproduced live in each panel's commentary strip, so a change to the CSS re-measures rather than re-asserting.
V0 is the only variant that loses words — 7 of 11 rows clipped at 240px, 11 of 11 at the 200px drag minimum. Its height never changes, because that is exactly what it is buying. Every other variant trades depth for the words back, and the question is only what the depth costs.
On the ikea set, just letting it wrap looks like the cheapest fix — and the Rockclimbing set says it is not enough. V1 costs +138px there and clips nothing, but on real Rockclimbing labels it clips 3 rows at 240px and 8 at 200px: the badge holds the first line, so the name gets roughly half the width for as long as it is sharing that line. Only the variants that give the name a line of its own — V2, V7, V9 — hold at zero across every set. That is a reversal of what this file said an hour earlier, and the real data is what reversed it. I also expected V4's regrouping to win, and it does not: at +159px it is taller than simply wrapping. The reason is visible in the panel — the Codebook-tags group has only one distinct badge value (“Sentiment”), so promoting it to a sub-heading removes six identical pills and adds a heading row, netting nothing. Grouping only pays where the column actually varies, which here is the Sentiment group and not the tag group. That is an argument for applying it per-group, not for the list as a whole.
V3 is nearly free on depth — −3px, because dropping the pill buys back most of what wrapping costs. It also reads best. It should not ship alone: colour as the sole carrier of sentiment fails a colour-vision read and needs a legend. It is here because rail + grouped sentiment is a real combination, where the sub-heading spells out what the rail merely echoes.
V2 and V5 both cost more than V1 for no additional words. V2 (+235px) is the most legible two-line row and the most expensive. V5's inline tail (+176px) forces a wrap wherever the badge doesn't fit after the last word, which on this content is most rows — it flatters short names and punishes long ones, which is the wrong way round for the set that has the problem.
Set the content picker to ikea · plain to see the other half of the bind: at 240px nothing truncates in any variant, so every one of these layouts is pure cost. The row is bimodal, and any layout chosen here is being chosen for the elaborated case while the plain case pays for it.
What six real projects say about V9
The second bench is not a stress test either. It is six screenshots: every
project here with enough analysis to fill a sidebar, rendered through the app's
own code path and cut by the same MAX_SIGNALS = 6. Six findings
came out of it, and not one of them is about width.
The Codebooks section is mostly the Sentiment section again.
In four of the six — rockclimbing, fossda, uxfriends, escuela —
every visible codebook row is the Sentiment group, so what
V9 draws beneath those names is the same closed set of seven words the section
above already spent its badges on. Counting rows rather than projects:
15 of the 34 visible codebook rows name a location that already has its
own row in Sentiment. escuela is 5 of 5.
And V9 is the variant that conceals it. A codebook row draws
signalName || location, so once elaboration has run the location is
not on screen at all: rockclimbing's “Dream Destinations and Travel
Aspirations” carrying delight in Sentiment, and its
“Destination Delight” carrying delight in Codebooks, are
the same cell four rows apart under two different names. The repeat count in each
commentary strip is computed from the data and not from the panel,
because the panel no longer contains it — which is the problem stated
exactly. Before elaboration the row shows the location and the duplication is
plain; elaboration is what hides it. That is a cost V8 does not have, since V8
keeps the location and changes only the badge.
Pairing can collapse a whole section to one row. ikea2's four
sentiment signals are all Beds Category — frustration,
delight, confusion, surprise — so V9 renders them as a single row carrying
four badges. That is the pairing doing exactly what it promised, and it is also
the ranking objection made visible: a ranked list of four became one row sitting
at the rank of its strongest member, with nothing on screen to say so.
The cap is a cliff, not a tax. Of those 34 rows, 8 carry more than three codes — and 6 of the 8 are fossda, the only project here with interview-length transcripts. Every other project shows every code it has at a cap of three. So the evidence that the cap is cheap comes entirely from the thin projects, and the one project with realistic data is the one where it binds on every row. Move the cap control and watch: only one panel changes.
A constant badge column is not only a codebook problem.
escuela's six sentiment rows are all frustration. V8's
argument — that one distinct value is a caption printed once per row rather
than a column — lands on the sentiment half too, on a project where it
happens to be true. V9 keeps that badge, so V9 inherits it.
One number above needs re-dating.
“11 of 11 projects have exactly one substantive codebook group” was
true of the corpus that measured it. Measured again on 12 Sep 2026 across the 25
project databases here that carry accepted tags, ten of which are the synthetic
stress-test set: of the fifteen real ones, eleven still carry exactly one
group, and four now carry two, three, six and nine. project-ikea prints
six different group names across its six visible rows, and ikea2 four. Those two
are also the only projects here whose V9 pills are open-vocabulary codes —
information architecture, did-it-work-ambiguity,
readme-as-first-resort — rather than sentiment words. So V8's
badge and V9's pills improve on exactly the same projects: the ones with more
than one codebook installed. That is a real argument for V9, and it is narrower
than “V9 shows the codes”: on eleven of fifteen projects there is
still one group, and the codes it shows are sentiments.
What this does not settle, and it is bigger than the row. The
six panels raise a question no variant on either bench answers: whether the
Sentiment group belongs in the Codebooks section at all, when it has
already been rendered above under its own heading. Excluding it is not a
cosmetic edit — fossda, uxfriends and escuela would have no
Codebooks section left, and rockclimbing would drop from six rows to two
(its only non-Sentiment rows are Feedback and
Status visibility, both below the cut). project-ikea and ikea2 lose
nothing; they have 23 and 12 non-Sentiment rows waiting. That is a question about
what the lens lists, not about how a row is laid out, and it changes what every
variant here is being judged on.
Why sentiment is in the Codebooks section at all
It looks like a blend of two kinds of data. It isn't. sentiment.yaml
is a codebook like any other — “Emotional & Cognitive
Signals”, one group, seven tags with definitions and
apply_when rules. At import,
_auto_tag_from_sentiment_field walks every quote and writes a
QuoteTag row for its Quote.sentiment value
(importer.py), so the codebook arrives pre-applied. From then on the
codebook analysis cannot tell it apart from a codebook you built by hand.
So the sentiment tags are not an ingredient in the other codebook
signals. Codebook analysis is partitioned by framework
(routes/analysis.py), and sentiment is its own framework — it
gets its own matrix, its own grand_total, its own signals. Nothing
about “Errors and recovery” changes if sentiment is removed. Hiding it
is a display decision, and it moves no other number.
What it does do is produce a one-column matrix, and that is where
the maths goes strange. concentration_ratio is
(cell / row_total) ÷ (col_total / grand_total). With one column,
col_total == grand_total, and every contribution in a row lands in that
single cell — so both halves are 1 and the ratio is exactly 1.00, in
all 45 Sentiment-group cells measured. Not roughly: structurally. Those
rows are ranked by a three-factor score with one factor switched off, against rows
whose same factor runs from 0.59 to 13.33. And
classify_flag is handed the group name, which isn't in
SENTIMENT_VALENCE, so they can never carry a finding flag either
— while the identical cells in the Sentiment section above can.
That is the sentence that is hard to say to a researcher, and the reason it is hard to say is that it isn't true of anything else in the lens: these rows are the same quotes as the section above, measured with a ruler that reads 1 every time.
The affordance already exists and Analysis is the one place that ignores
it. The codebook's own preamble tells the researcher to
“hide this layer with the eye toggle as you develop your own meaningful
tags”, and TagSidebar has that toggle. It writes
hiddenTagGroups into SidebarStore, the Quotes lens honours
it — and AnalysisPage.tsx never reads it. So this may not need a
new rule so much as one lens catching up with the rest.
One flat list
The fourth treatment drops the headings entirely: no Sentiment / Codebooks split, no Section / Theme sub-heading, no six-per-section cut. Every signal the project produced, sentiment rows paired by location as V9 pairs them, codebook rows as V9 pill rows, ranked strongest first and cut at a floor. Three things it shows.
Composite becomes comparable once the Sentiment group is gone. The degenerate one-column matrix was the whole problem. On the two projects with real codebook data the remaining ranges genuinely overlap — ikea2 runs 0.92–1.85 on sentiment against 0.80–2.25 on codebooks; ikea 2.00 against 0.77–3.86. Interleaving those is defensible in a way that interleaving the shipped three populations was not.
An absolute floor is not portable between projects. At 0.40, rockclimbing keeps 11 of 14 and ikea2 keeps 3 of 13 — not because one project has three times the insight but because composite is not normalised across studies. A floor expressed as a share of that project's strongest signal would behave; this one is the instrument, not the proposal.
Without the group heading, repeated locations read as duplicates. project-ikea has one location carrying eight separate codebook rows, and 12 of its 23 rows are un-elaborated — only the top of each partition gets a signal name, so the rest draw the bare location. In the flat list that is eight rows saying “Beds & Mattresses Category”, told apart only by their pills. The two-section layout hid this behind a heading; the flat list has to answer it, either by elaborating every row or by pairing codebook rows on location the way the sentiment half already pairs.
Worth saying plainly: on four of these six projects the flat list is the sentiment list, because those projects have no codebook tags that aren't sentiments. The flat list does not create insight that the study did not produce. What it does is stop the lens from presenting one study's worth of findings as two.
The codebook pills had no colour, and that was hiding V9's best argument
A bug first. Every panel above referenced
--bn-ux-1-bg and its siblings and this file never declared them, so
each codebook pill silently fell back to the plain grey badge background. All of
the V9 panels have been arguing for showing the codes while drawing them in the one
colour that says nothing about which codebook they came from. The 26 slot tokens
are now inlined verbatim from theme/colors/palette-*.css.
And the mechanism is per-code, not per-group.
getTagBg(colourSet, index) resolves
var(--bn-<set>-<(index % slots)+1>-bg), and
_build_tag_colour_indices assigns that index in tag-definition order
within each group. So two codes in one group are two tints of one hue, and
two groups are two hues. The panels now carry the real indices, read from each
project's database, not a single tint per group.
It buys less than I first claimed, and the five-codebook project is what
showed that. An earlier version of this paragraph said the hue carries
“which codebook group this row belongs to”. It does not.
colour_set is a property of the group, and there are only six of them
for any number of groups. project-ikea has fourteen groups sharing six
colour sets across four codebooks: ux covers Discoverability,
Strategy, Behaviour, Real-world matching and Status visibility; task
covers Structure and Conceptual model; emo covers Feedback, Scope and
Needs and desires. On its busiest section, four chips would be the same blue and two
the same green. The hue is a family cue, not a label, and the group name
cannot leave the row on its strength.
What the hue does buy is real but narrower: a code keeps the same tint everywhere,
so did-it-work-ambiguity is trust-4 in every row it appears
in and is recognisable without being read; and two codes in one group read as
obviously related. That is worth having. It is not a replacement for saying which
group a row belongs to.
One collision to know about.
--bn-ux-5-bg and --bn-trust-1-bg are both
oklch(94% 0.03 275) in light mode — the same hue angle, the same
colour. ikea2's flat list has search-engine-fallback (ux, slot 5) and
error-message-illegible (trust, slot 1) rendering identically. With
five groups on screen at once, the palette's wrap-around stops being theoretical.
Should a codebook pill be clickable?
First, what is actually clickable today. Nothing, individually.
The shipped SignalEntry is one <a> wrapping the name
and the badge together — the whole row is one target. Per-badge links are a
proposal, invented here by V6 / V7 / V9's paired rows. So this is not an
existing asymmetry to extend; it is a new affordance, and the question is whether
its counterpart should exist at all.
The answer falls out of the partition / relation distinction above. A sentiment badge in a paired row is a signal: (location × sentiment) is a real cell with its own card, and pairing folded several of those rows into one. Each badge has to stay clickable or those signals become unreachable — the link is recovering what the pairing took away. A codebook pill is not a signal. The signal is (location × group); the pills describe what is inside that one cell. There is no card at (location × tag) to go to.
So two rows that look identical mean opposite things: one row's pills are a list of destinations, the other row's pills are a description of one destination. In the two-section layout a heading warned you which you were looking at. In the flat list nothing does.
Four ways out, and they are not equally good:
- Leave them inert. Cheapest and worst in the flat list — two identical controls, one live, one dead, with nothing to tell them apart.
- Filter the card's quotes by that code. Honest: “show me
the quotes in this signal carrying this one”. The destination already exists
— the card lists the quotes and each carries
tagNames— and it does not claim the pill is a signal. My pick for the click. - Show the code's definition. The vocabulary is open and often
not the researcher's own;
did-it-work-ambiguityearns a definition. The codebook has one, withapply_whenandnot_this. That is a hover, not a click — the pills now carrygroup · codeas a title attribute as a first cut. - Run the codebook analysis at tag level, columns = tags rather
than groups. Then each pill genuinely is a signal and the symmetry is real rather
than cosmetic.
detect_signals_genericwas built for exactly this — its docstring offers “codebook groups, individual tags, or any other categorical dimension”. Measured cost: across the three projects with non-sentiment codebook data, 37 group-level signals become 61 tag-level cells, of which 22 clear the existing 2-quote floor (ikea 13, ikea2 9, rockclimbing 0). It also moves concentration again — more columns means a smaller expected value means a higher ratio — so it reopens the comparability question the flat list had just closed.
Recommendation: filter on click, define on hover, and decide tag-level
analysis on its own merits — whether a signal at
(location × one code) is a thing a researcher wants, not whether it would make
the pills clickable. And whatever the answer, a live badge and an inert pill should
stop being the same control. That is a design-system question this bench cannot
settle by measuring, because both are currently .badge.
Normalising two analyses into one ranking
Interleaving needs the two scores to mean the same thing, and they half do.
composite = conc × (n_eff / participants) × (intensity / 3).
The last two are matrix-independent: same quotes, same arithmetic, both 0–1.
Only conc carries the shape of the table it came from.
The good news is that a lift ratio is already centred.
conc is observed ÷ expected, so its value under independence is
1.0 in any table of any width. Its centre needs no normalising.
The bad news is the ceiling. The most a lift can reach is
grand_total / col_total, and that is pure table shape. Measured across
these six projects:
sentiment matrix 7 columns ceiling 1.46 – 53.33 (median 3.40)
codebook · other 3–10 columns ceiling 1.00 – 7.33 (median 3.00)
codebook · Sentiment 1 column ceiling 1.00 – 1.00
So a lift of 2.5 is nearly the maximum in one table and unremarkable in another, and the same number ranks two signals that are not equally surprising. The 53× tail is a rare sentiment in a small project — which is exactly the row that will top a merged list for the wrong reason.
What the flat list above actually does is nothing. It ranks on raw
composite with the Sentiment group removed, and gets away with it
because the two median ceilings land at 3.40 and 3.00. That is luck, not design, and
it is stated here rather than buried so the next person does not inherit it as a
decision.
The statistic that fixes it is already in the tree and ranks nothing.
adjusted_residual in metrics.py is the standardised
residual — the textbook answer to comparing cells across differently-shaped
contingency tables, and currently used only to colour heatmap cells. Measured on
the same signals:
sentiment matrix z −1.37 … 6.08 median 1.47 |z|>2 on 12 of 36
codebook · other z −0.74 … 2.48 median 1.19 |z|>2 on 6 of 37
codebook · Sentiment z 0.00 … 0.00 — the same structural fact, said properly
Medians 1.47 against 1.19, where the lift ceilings were 3.40 against 3.00 with a
53× tail. Much closer, and |z| > 2 means the same thing in
both. It also reprices the list honestly: 18 of 118 signals across six projects are
notable by that test.
But it cannot simply be dropped into the composite, and the reason is a
product decision rather than a maths one. The composite multiplies,
and z goes negative — a negative residual means this place has notably
fewer of these than expected. Multiply that by breadth and intensity and a
strong absence sorts to the bottom of the list, when a strong absence is a finding:
nobody mentioned price on the checkout screen. So adopting z forces a choice —
rank on |z| and show direction, clamp at zero and discard depletion, or
keep two lists. Worth noticing that classify_flag already has this
vocabulary (Win, Problem, Niggle, Success, Surprising) and already gates nothing.
Several signals about one place is the finding, not the duplication
The flat list made a location's repeats look like noise — eight rows saying “Beds & Mattresses Category”. Read the other way round, that is the most useful thing the analysis found: the checkout is empowering and confusing; pricing reads expensive and fair, and separately shocks on a change and delights on a discount. A navigation that flattens those into one row per place loses the findings; one that lists them separately loses the place.
The last two treatments keep both: one entry per location, every signal under it its own destination. This is not a new mechanism — it is V7's pairing with the section boundary removed. V7 already does exactly this for the sentiment half and stops there.
Measured on ikea: 8 places, 18 destinations, 5 of the places carrying more
than one signal, 5 on the busiest. The busiest reads
confusion · Behaviour ·
Structure · Conceptual model ·
Real-world matching — one sentiment and four codebook groups
about the same category, each reachable. ikea2's Beds Category reads
frustration · delight ·
confusion · Errors and recovery, which is the
shape of the problem in one line.
And it gives the codebook chip an honest click. The objection above was that a codebook pill is not a signal — the signal is (location × group), and a code is only a description of it. Group by location and the chip becomes the group, which is exactly a signal, with a card to land on. The tension dissolves: the thing that could not be clicked was a code; the thing that can be is a group.
The cost is the codes themselves. The compact variant spends its badges on destinations, so V9's whole argument — show the codes actually present — is gone. The named variant buys them back at a sub-row each, which is the depth V9 was trying to avoid spending. Nothing here settles that; it is the next thing to look at rather than reason about.
What the five-codebook project settles
Five of the six projects on the second bench cannot tell these treatments apart: they have no codebook tags that are not sentiments, so every treatment draws them the same. project-ikea is the only real test — four codebooks, fourteen groups, 23 non-Sentiment signals, eight locations. The third bench runs it through every treatment at once.
The sentiment floor is a volume floor, and on a small study it empties the
sentiment half entirely. MIN_QUOTES_PER_CELL = 2 in
signals.py asks for two quotes in a cell, not for a strong one.
project-ikea clears it once. Lower it to 1 and the same analysis
returns 18 sentiment signals across 10 locations — real
cells, real composites, one quote each. That is the
sentiment volume control, and it is why this project looked like a
codebook-only study: not because its participants said nothing with feeling, but
because three participants rarely put two quotes in the same cell.
And those rows break the ranking, which is the normalisation problem
arriving in person. At ≥1 the flat list is topped by
“Duvets & Bedding Category · surprise” and
“App Prompt & QR Code · surprise” — single quotes from
single participants, scoring 1.19 because a rare sentiment in a
small study has a lift ceiling of 5.33. Every codebook signal in the project is
below them; the best is 0.64. One person being surprised once outranks eight
codebook findings. That is the strongest case yet for the standardised residual, and
it is not visible on any project that has only the shipped floor's survivors.
One section carries ten signals, from four codebooks. Beds &
Mattresses Category holds confusion and frustration from
the sentiment matrix, plus Behaviour, Structure, Conceptual model, Real-world
matching, Status visibility, Feedback, Needs and desires and Discoverability —
ranked 0.67 down to 0.11. So yes: a section or theme routinely has several signal
cards from several codebooks, and each has to be its own target. It is also true
that the busiest place is busiest mostly with weak signals, which is what
makes the strength floor the instrument that matters rather than the row layout.
Measured on this project, grouping by location is the denser navigation by a wide margin:
as shipped 12 rows 12 targets 20 chips
hide + backfill 12 rows 12 targets 19 chips
one flat list 27 rows 28 targets 47 chips
by location 12 rows 43 targets 31 chips
by location · named 12 rows 43 targets 43 chips
Same twelve rows as the shipped layout, three and a half times the destinations. The flat list needs more than twice the depth to reach two-thirds as many. That is the pairing argument generalised: a location appears once, and everything known about it hangs off that one entry.
Does a bigger study fix the empty sentiment half?
The obvious objection to the volume-floor finding is that these are trial projects.
A real study is five sessions of fifty minutes, not three of a few — so cells
would commonly hold four or five quotes and MIN_QUOTES_PER_CELL = 2
would never bind. Measured across every analysed project here, counting quotes that
carry both a sentiment and a location, which is what the sentiment matrix
actually consumes:
sess quotes q/sess locs cells median mean ≥2
ikea 3 20 6.7 10 18 1 1.1 1/18
ikea2 3 27 9.0 8 19 1 1.4 4/19
rockclimbing 8 81 10.1 23 45 1 1.8 17/45
fossda 20 167 8.3 17 58 2 2.9 38/58
Quotes per session barely moves, and it is nearly independent of session length. 6.7 to 10.1 across projects whose sessions run from about three minutes to about fifty. fossda's are roughly 50 minutes each and yield 8.3 sentimented quotes; ikea's are minutes long and yield 6.7. A longer interview does not produce proportionally more quotes — extraction picks the notable moments, and that count is close to bounded per session.
Occupancy does rise with study size, but sub-linearly, because locations multiply alongside quotes: 3 sessions gives a mean of 1.1–1.4, 8 gives 1.8, 20 gives 2.9. Interpolating, a five-session study lands near a mean of 1.5 and a median of 1 — not four or five. Even at twenty sessions and sixteen hours of material, the median populated cell holds two quotes, and only 15 of 58 reach four.
And the binding constraint is not quotes at all. Median
distinct participants per populated cell is 1 in every project,
fossda included. A cell with four quotes is usually one person four times
— which is exactly what simpsons_neff measures and what
confidence gates on, and it is the right thing to be scarce.
So the volume floor softens with scale rather than disappearing, and this
is a prediction to check rather than a conclusion to keep. Richer test data
is coming. When it lands, the numbers to re-measure are: mean quotes per populated
sentiment cell at five sessions (predicted ~1.5, median 1), the share of cells
clearing MIN_QUOTES_PER_CELL (predicted roughly a quarter), and median
participants per cell (predicted 1). If the real figure is a median of four
or five, this whole section is wrong and the volume-floor finding goes away with
it — which is the point of writing the prediction down.
The section above measures the wrong axis
Everything on this page about thin data measured the sentiment matrix, and
sentiment is bounded at one value per quote by construction —
Quote.sentiment is a single dominant value or None. It can never be
dense. On ikea it runs 0.61 sentiments per quote: two in five quotes
carry none at all. That is not a shortage of material, it is what a partition over a
closed set of seven looks like. Interviews are not all high emotion, and a study that
produced a fat sentiment matrix would be a strange study.
Codebooks are a relation, and that is the whole point of them. A quote carries as many codes as apply, from as many groups as are installed. Measured on the same quotes:
quotes sentiments/quote codes/quote groups with data
ikea 33 0.61 3.73 23
ikea2 38 0.71 1.21 9
uxfriends 28 0.54 1.25 12
rockclimbing 78 0.82 0.13 3
Sentiments per quote barely move — 0.54 to 0.82 — because that number is a property of the extractor, not of the study. Codes per quote move by a factor of 29, and they move with exactly one thing: how many codebooks were applied. So the lever on data volume is not more interviews, it is more codebooks — and a domain expert's own codebook adds a whole dimension of coverage to quotes that already exist. rockclimbing and fossda look sentiment-only here because nobody applied a codebook to them, not because their material was thin.
Which reprices the earlier calibration. The session-count analysis above is still correct about the sentiment matrix and still predicts the same numbers, but it was answering a question that does not decide much. The sentiment half is meant to be sparse — the codebook's own preamble calls it a layer to orient by and then hide. Whether the lens has enough to show is a question about the codebook half, and that half is 6× denser on the one project here where codebooks were actually used.
One thing a reader should know about that density. ikea's 3.73
codes per quote is 31 accepted tags and 112 pending proposals:
_compute_group_analysis includes pending ProposedTag rows
weighted by the model's confidence, so most of what the Codebooks section is ranking
today is AutoCode's unreviewed suggestions. That is the design working as intended
— propose, then let the researcher accept — but it means the signal
strengths in that half are partly confidence-weighted guesses, and they will move as
a researcher works through the queue.
More codebooks is not more coverage — and that is the stress test
The section above said the lever on volume is more codebooks. Half right, and the half that is wrong matters more. Measured on ikea's four UX codebooks, pairwise overlap on the quotes they tag:
uxr × nielsen 1.00 — the same 33 quotes, exactly
uxr × norman 0.91
nielsen × norman 0.91
uxr × garrett 0.82
norman × garrett 0.78
25 of the 33 coded quotes carry codes from all four codebooks. So codes per quote rises with each one installed — 3.73 here — while the set of quotes covered does not move at all. A second UX codebook is not finding material the first missed; it is re-describing the same material in another vocabulary. Coverage saturates at one codebook; what grows is the number of ways each quote is said.
Which reprices the budget. ikea's 23 non-Sentiment signals span eight locations. A list of 25 findings drawn from four overlapping frameworks over 33 quotes is not 25 findings; it is closer to six or eight described several times. Under one location the near-synonyms sit adjacent: Feedback (Norman) beside Status visibility (Nielsen), Conceptual model (Norman) beside Real-world matching (Nielsen). The two-section layout kept them apart by accident. Grouping by location puts them side by side, which is honest and is also the hardest thing this layout has to survive.
So ikea is the stress test, not the target configuration. The expected shape is one sentiment codebook, one or two UX expert codebooks, and later a client- or project-specific one — custom codebooks are not a day-one offer. At one or two UX codebooks a quote carries one or two descriptions, and the adjacency problem is mild. Five is the upper bound the information design has to hold at, and it is the right thing to draw against precisely because it is worse than what ships.
One budget over one ranking
The deliverable is 10–25 findings from five to twelve interviews. Against that,
MAX_SIGNALS = 6 applied twice is the wrong shape: it caps the lens at
twelve and, more to the point, it under-fills whenever one side is thin,
which is most projects. Slots delivered against a 12-slot allowance, with the
Sentiment group excluded:
sentiment codebook 6+6 one 12 wasted
rockclimbing 12 2 8 12 4
ikea 1 23 7 12 5
ikea2 4 12 10 12 2
fossda 12 0 6 12 6
Four of six projects leave slots empty, fossda half of them. The
signal budget control replaces the two caps with one number over one
ranking: floor first (is this worth showing), budget second (how many will anyone
read). Each commentary strip now says which one bound the list —
budget, floor or supply — because those
are three different problems and only the last one needs more data.
Note what the budget does not settle, and it is the thing the overlap measurement just made urgent: filling 25 slots is easy, and filling them with 25 distinct findings is not. Nothing in the ranking knows that Feedback and Status visibility are the same observation in two vocabularies.
Lifting a winner into the app
Each variant shares one markup contract, so the shipped component only has to grow a class:
SignalEntryalready exists as a discrete component inAnalysisSidebar.tsx— extract it to its own file and give it avariantprop.- V1, V2, V5 are pure CSS on the existing markup.
- V3 needs a
--railcustom property from the sentiment name, whichgetGroupBg()already does for the tag colour sets. - V4 needs the grouping moved up:
signalsBySourceType()gains a sibling that buckets bycolumnLabel, and the badge stops rendering per row — worth doing only for groups where that value varies.
Whatever wins, turn “break long words” on and check the torture set.
The shipped .signal-entry-name sets no overflow-wrap, which
costs nothing while the name is ellipsed on one line and starts mattering the moment
it is allowed to wrap. An unbreakable token — a compound identifier, a URL, a German
noun — overflows the panel silently in every wrapping variant until that property is
set.
The width slider stays useful after a variant wins — it's the check that the chosen layout still holds at the 200px drag minimum, which is where a user who wants their content back will put it.