AI-Assisted Product Design · Case Study CompanionThreadwise v0.1
One frozen brief,
three model minds.
The same Threadwise foundation went to three Claude Design models — Sonnet, Opus, and Fable — each running the same staged process: propose three directions, resolve one, build hi-fi screens, then a working prototype. This is the evaluation: what held, what drifted, and where a designer's judgment still had to close the gap.
Most complete and the only prototype with a live scoring engine — including the multi-component case none of the others wired.
The cleanest, most consistent execution and the tightest scope guardrails — but the thinnest coverage of edge cases.
The strongest design thinking and the most principled honesty — and the three directions the final hybrid was actually built from.
None of the three broke scope. No model added comparison, accounts, a wardrobe, streaks, or "buy this instead." Every difference below is about craft, completeness, and judgment — not hallucination. Against a well-frozen brief, capable models behave. The designer's real work moves up the stack.
Executive summary
Three models received identical Threadwise foundation documents and a near-identical staged prompt sequence. The experiment tests a simple, useful question for anyone using AI in product design: given a tight, well-specified brief, do capable models stay inside it — and where do they still need a human?
The reassuring headline is that all three respected the frozen MVP. Each kept Threadwise a single-garment tool, each explicitly excluded the Phase 2 comparison flow, and none invented features, accounts, or purchase suggestions. Two of the three even restated the guardrails on-screen ("single garment," "Phase 2 excluded," "no suggestions") before designing a pixel.
The differences showed up one layer down. Sonnet optimized for completeness and working systems: 18 fully specified screens and the only prototype with a real parser and scoring engine, including live multi-component handling. Opus optimized for restraint: the calmest, most internally consistent visual system and the strictest scope discipline, at the cost of several edge states it simply didn't draw. Fable optimized for design judgment: the richest exploration, an explicit "held constant" frame, per-direction interaction specs, and the most principled honesty behaviour — it refuses to invent a score when a label is only partly readable.
Model choice here is a choice of strength, not correctness. Fable is the strongest single body of design thinking (and the directions the final hybrid was built on); Sonnet is the strongest proof of functional execution; Opus is the strongest demonstration of restraint. The most honest recommendation is to use all three as evidence of different senior skills — with the human synthesis on top as the real payoff.
Experiment setup
A controlled comparison: same inputs, same process, three model "brains." The only variable that changed was the model doing the work.
| Stage | Ask | What it tests |
|---|---|---|
| 01 Three directions | Propose three distinct hi-fi visual directions for the frozen MVP. | Range of design thinking; can it stay in scope while exploring? |
| 02 Resolve one | Select and commit to a single direction (or a deliberate hybrid). | Editorial judgment; can it choose and justify? |
| 03 Hi-fi screens | Build every MVP screen and state in the chosen direction. | Completeness; edge-case coverage; design-system fidelity. |
| 04 Prototype | Wire an interactive, tap-through version of the flow. | Interaction logic; does the pattern actually work when clicked? |
The brief states the prompts were "nearly the same" across models. The prompts themselves were not part of this review — only the nine output artifacts and their underlying code. Where a model's stage intent is inferred, it is drawn directly from that artifact's own labelling (each model explicitly named its stage and directions). Scores reflect the evidence in the artifacts, spot-checked against the source documents; they are a designer's calibrated judgment, not an absolute metric.
What the foundation controlled
The Threadwise pack froze six things before any model touched it. These are the yardsticks every output is measured against — and the source of the one genuine ambiguity the models had to navigate.
| Control | What it locked |
|---|---|
| MVP scope | Text-first, single garment, sustainability-led, with a durability-and-care layer. Comparison of 2–5 garments is Phase 2, explicitly out of the MVP. No accounts, wardrobe, brand ratings, notifications, or feeds. |
| Scoring logic | A 0–100 Fabric Composition Score from a fixed fiber base-score table (linen 90 → elastane 20), weighted by percentage, with two small blend penalties (−5 mixed, −5 complex). An "illustrative, meant to guide" cue and a permanent "what this can't see" line. |
| Design system | The "festival, not a sacrifice" heritage system: the jharokha arch motif, Heritage Pink #8B2149, cream #FBF4E6, Rozha One display serif over Hind body over IBM Plex Mono. "Green is earned," one accent per screen, never colour alone. |
| Hi-fi behaviour | The result is a card that rises over a dimmed Home. An eight-zone anatomy: verdict + score, honesty cue, one reason line (peek ends here), breakdown, end-of-life/microplastic, durability/care, an inline "why this score," and one calm action. |
| Interaction rules | Peek → expand → "why," all tap-driven, no drag physics. One overlay card maximum. "Why" is an inline accordion, never a second sheet. About, Error, and Compare are full pages — never dressed as result cards. Degrades to a single panel on desktop. |
| Content tone | Calm, plain, warm, non-judgmental. "Better to skip," never "avoid" or "bad." "Breaks down," not "biodegrades." State the verdict and step back — no upsell, no guilt. |
The foundation is not perfectly self-consistent, and this matters for judging the models fairly. The Source of Truth and Hi-Fi Spec describe three calm bands (Good to buy / Okay / Better to skip) and a "why" that opens five lifecycle dimensions. The dedicated Fabric Composition Score document narrows the MVP to composition only, uses five bands, a three-line explanation, and parks the lifecycle dimensions for v2. No model created this tension — but every model had to resolve it, and they resolved it differently. Watch it recur in the drift analysis.
Model-by-model analysis
Each model, read across all four stages and its underlying code. Same rubric, three temperaments: the engineer, the minimalist, and the design lead.
Sonnet · the engineer
Sonnet treated the brief as a system to be fully built. Its three directions vary by density and craft — Minimal & fast, Score-card focused, Guided explanation — and it resolved them into a single "Plan A" hybrid, then shipped the most complete artifact set of the three.
Where it followed the source well
- Embedded the frozen fiber base-score table verbatim (linen 90 … elastane 20) with synonym handling (rayon, polyamide, lycra).
- Handled all six named edge cases, including the multi-component "verdict reflects the shell, lining shown separately" rule.
- Held the card pattern precisely: tap-driven peek → expand → inline "why," About and errors as full pages.
- Tone stayed calm and factual — the Skip variant reads "…working against you on fabric alone," never "bad."
Where it drifted
- Used the five-band scale in its directions but the three-band scale in its hi-fi — an internal naming inconsistency to reconcile.
- Showed the five lifecycle "why" bars, which the composition-only scoring doc parks for v2 — mitigated by a "not five separate scores" caption.
- By its own admission the plainest direction "could be any utility app" — the least distinctive on brand of the three.
Functional completeness. It is the only submission where you can type "70% wool, 30% polyester / lining 100% polyester" and get a correct, live, multi-component verdict — the clearest proof the product's logic actually works.
Visual distinctiveness. The resolved direction is clean and legible but the least characterful; it spends its effort on coverage and correctness rather than making Threadwise feel unmistakable.
Human correction required
Pick one band scheme and propagate it; decide whether the "why" should expose five dimensions or hold to composition-only; push one notch further on brand expression if this direction is taken.
Opus · the minimalist
Opus treated the brief as something to execute with restraint. Its three directions — The Doorway, The Care Label, Calm Focus — are the most confidently art-directed, and it carried the frozen guardrails onto the canvas as literal chips before designing a thing.
Where it followed the source well
- Most internally consistent band model — three calm bands with the Source-of-Truth ranges (70–100 / 45–69 / 0–44), held across every stage.
- Used the trace-synthetic honesty line almost verbatim ("…can't be recycled as a natural fiber") and the labelled fact rows (END OF LIFE / MICROPLASTIC / ON SKIN / CHECK YOURSELF).
- Named the card-pattern boundary out loud: "Error and About are full screens, never result cards."
- Highest design-token fidelity — leaned fully into the arch and heritage system without inventing off-brand colour.
Where it drifted
- Drift by omission: several edge states the spec calls its "strongest-specified area" — multi-component, partial read, invalid input, permission-denied, low-light — are not drawn or wired.
- Also showed the five "why" bars — the same tension with the composition-only scoring doc.
- The Doorway's ceremony carries a cost its own notes flag: ~90px of arch header before any facts, which strains on rapid rack re-checks.
Visual discipline and consistency. The Doorway is the most polished single direction, and the band-and-token system is the most coherent of the three — the best evidence of taste and restraint.
Completeness. It shows the product at its best but not at its messiest; a designer adopting it inherits the job of drawing every unhappy path the others already covered.
Human correction required
Backfill the missing edge states (multi-component, partial, permission, low-light); resolve the "why" content against the frozen scoring model; pressure-test the arch header for the fast aisle case.
Fable · the design lead
Fable treated the brief as a design problem to reason about out loud. It shipped the richest exploration — a shared "frame every direction walks," per-direction interaction specs, and a closing synthesis that reads each direction as a strategic bet.
Where it followed the source well
- Most faithful to the scoring doc's language — used its exact recommendation labels ("solid pick," "weak fabric choice," "probably skip").
- Best embodiment of the honesty principle: on a partly-unreadable label it shows "Can't fully score … the rest is a guess we won't make. No score is better than a made-up one." — it refuses to invent a number.
- Locked the frozen words explicitly ("fiber" not "fibre," calm verdicts never "avoid," cue rides with the score, no scoring math in the flow) as a stated constraint.
- Anticipated the exact hybridization path — "the Signal's peek with the Docket's expanded rows" — before a direction was even chosen.
Where it drifted
- Declared "no scoring math in the user-facing flow" in its directions, then reintroduced the five "why" bars in its hi-fi — the same latent tension, and a small inconsistency with its own rule.
- Blends three-band headline words with five-band sublabels, so some screens carry two verdict labels at once — thoughtful but slightly redundant.
- Its prototype wires partial and unrecognized states but not the multi-component case Sonnet handles live.
Design judgment and narrative. It doesn't just produce screens, it produces an argument — and its honesty behaviour (refusing a made-up score) is the single most on-brand decision in the whole experiment.
Discipline in its own detail. The double-labelling and the directions-vs-hi-fi "why" reversal show it occasionally out-reasoning its own constraints; it needs a light editorial pass to prune.
Human correction required
Resolve the "why" content once (its directions had the more principled instinct); prune to a single verdict label per screen; add the multi-component case to match its own thoroughness elsewhere.
Comparison matrix
All twelve criteria at a glance. The bar shows strength; the colour shows which model. Read a row to compare the field on one dimension, or a column to read one model's shape.
| Criterion | Sonnet | Opus | Fable |
|---|---|---|---|
| MVP scope control01 | Strong Single garment held; "no comparison, no suggestions" stated up front. |
Strong "Phase 2 excluded" carried onto the canvas as a literal guardrail chip. |
Strong Froze scope + words in a dedicated "held constant" frame. |
| Product clarity02 | Strong "A huge score, one sentence, done" — unmissable value. |
Strong "Answer in one breath, then depth on tap" reads instantly. |
Strong Poster-scale question framing; the answer legible before a word is read. |
| In-store usability03 | Strong Fastest read at every step; quiet always-available exit. |
Solid Beautiful, but the arch ceremony adds overhead on rapid re-checks. |
Strong The Signal is built for the two-second, one-handed aisle glance. |
| Result-card fidelity04 | Strong Peek → expand → inline "why," one overlay, tap-driven throughout. |
Strong Named the boundary: About/Error as full pages, never cards. |
Strong Pinned verdict head, inline "why," scrim-to-dismiss — spec-exact. |
| Visual design quality05 | Solid Clean and legible, but self-described as the least "branded." |
Strong The most polished, considered surfaces of the three. |
Strong Three fully-realised design arguments, each unmistakable. |
| Design-system alignment06 | Strong Correct tokens + font stack; no off-brand colour. |
Strong Highest token fidelity; fully committed to the arch system. |
Strong Even matched the documentation pack's own chrome. |
| Content tone07 | Strong Calm and factual; Skip never reads as blame. |
Strong Honesty cue + "check yourself" line held consistently. |
Strong Warmest, most human copy; "no score is better than a made-up one." |
| Interaction quality08 | Strong Wired the most states, including live edge cases + dev triggers. |
Solid Cleanest happy-path wiring; fewer states reachable. |
Strong Best-specified motion (timing, scrim, reduced-motion) and wired refusal. |
| Edge-case coverage09 | Strong All six edges drawn and wired, incl. live multi-component. |
Partial Missing multi-component, partial, invalid, permission, low-light. |
Strong All edges + extras; most principled partial-read handling. |
| AI drift resistance10 · higher = less drift | Solid No invention; band-scheme inconsistency + five-bar "why." |
Strong Most consistent; drift is by omission, not invention. |
Solid No invention; out-reasoned its own "no math" rule once. |
| Low human correction11 · higher = less needed | Solid Reconcile bands; decide "why" content; lift brand a notch. |
Partial Inherits the most backfill — every missing edge state. |
Solid Prune double-labels; settle "why"; add multi-component. |
| Portfolio strength12 | Strong Proof the system works — the execution exhibit. |
Solid Proof of restraint — the art-direction exhibit. |
Strong Proof of judgment — the design-thinking exhibit, and the hybrid's source. |
| Overall score | 8.5 / 10 | 8.0 / 10 | 9.0 / 10 |
Key findings
What this experiment revealed about using AI as a design material — beyond the three individual scorecards.
Scope held everywhere. Against a tightly frozen brief, all three models stayed inside it — no comparison flow, no accounts, no wardrobe, no gamification, no upsell. The classic fear ("the AI will run off and redesign my product") did not materialise once. A well-written brief is the most effective guardrail there is.
02 · Capability = completeness + judgment
The gap between models was never the happy path — all three nailed that. It was coverage of the hard 20% (edge cases) and the quality of reasoning about tradeoffs. That's where "strong" pulled away from "adequate."
03 · The brief's own gaps propagate
Where the foundation contradicted itself (three bands vs five; composition-only vs five dimensions), each model faithfully followed one source doc. Faithfulness to a conflicting spec looks like drift but isn't — it's a spec problem only a human can settle.
04 · Working prototypes are table stakes
None of the three shipped a clickable mockup. All built real parsers and scoring engines from the written fiber table — Sonnet even computes multi-component verdicts live. The bar for "prototype" has moved from "looks interactive" to "is interactive."
05 · Edge cases are the tell
Same peek-expand-why on the happy path; wildly different unhappy paths. If you want to tell a strong model output from a passable one quickly, look at what it does when the label is torn, blurry, unnamed, or lined.
Fable's refusal to invent a score when a label is only partly readable — "no score is better than a made-up one" — is a values decision, not a UI decision. It's the hardest thing to specify in a prompt and the most worth noticing when a model does it unasked. That instinct, more than any layout, is what "on-brand" actually means for an honest-broker product.
AI drift patterns
Where the models overreached, over-explained, or drifted visually. The honest read: drift here was mild, and mostly not hallucination — it was faithfulness to a brief that quietly disagreed with itself.
| Pattern | Where it showed up | Severity |
|---|---|---|
| Five-dimension "why" | All three, in hi-fi. The "why this score" opens five lifecycle bars (Climate/Water/Land/End of life/On skin) — faithful to the Hi-Fi Spec, but the composition-only scoring doc parks these for v2. All three softened it with a "not five separate scores" caption. | Low |
| Band-scheme divergence | Sonnet used five bands then three; Fable blended three headline words with five sublabels; Opus stayed three throughout. No model invented bands — they chose among the brief's own two schemes. | Low |
| Out-reasoning its own rule | Fable declared "no scoring math in the user-facing flow," then reintroduced the five "why" bars in hi-fi. A small self-inconsistency from a model thinking harder than its own constraint. | Low |
| Verbosity / double labels | Fable carries two verdict labels on some screens (e.g. "PROBABLY SKIP" + "weak fabric choice"). Reads as thoroughness, but adds words a glance-first product doesn't need. | Low |
| Drift by omission | Opus left several spec-emphasised edge states undrawn (multi-component, partial, invalid input, permission, low-light). The most consequential drift here — invisible until you audit for what's absent rather than what's wrong. | Medium |
No invented features. No sneaking in the comparison flow. No "buy this instead." No raw scoring math dumped on the user. No off-brand colour. The failure modes people most associate with AI-generated design were absent across all three — the drift that remained was subtle and interpretive, not reckless.
Where human judgment was necessary
The models widened the option space and built fast. Narrowing it, reconciling the brief's contradictions, and making the taste-and-values calls stayed human work. This is the real division of labour the case study demonstrates.
- Reconcile the band scheme. Three bands or five? The brief offers both. A designer has to pick one and propagate it everywhere — a product decision no model could make on your behalf.
- Decide the "why" content. Five lifecycle bars (per the Hi-Fi Spec) or composition-only (per the scoring doc)? The models defaulted to five; the frozen scoring model argues for restraint. This is a judgment call about how much logic to expose.
- Choose and hybridise a direction. The final architecture — The Doorway's structure, The Docket's data rows, The Signal's interaction rules — is a human synthesis across Fable's three directions. No model performed this cross-pollination; Fable only predicted it was possible.
- Backfill the gaps. If Opus's visual direction wins, someone has to draw every unhappy path it skipped — the multi-component, partial, permission, and low-light states.
- Prune the over-reasoning. Hold Fable to its own "no scoring math" instinct and cut the double-labelling; keep the thinking, lose the redundancy.
- Make the taste-and-values calls. Which honesty behaviour to keep (Fable's refusal to guess), how much ceremony the aisle can bear (Opus's arch), how plain is too plain (Sonnet's minimal). None of these has a "correct" answer the model could look up.
AI moved the designer's work up the stack: less pixel-pushing, more direction, critique, reconciliation, and judgment. The skill on display in this case study is not "made screens with AI" — it's directed three models, read their output critically, and added the judgment they couldn't.
Recommendation
If the portfolio needs a single "strongest output," it's Fable — narrowly. But the more honest and more impressive move is to use the three together, each as evidence of a different senior competency.
Lead with · Fable
For depth of design thinking, the most principled honesty behaviour, and the fact that the final hybrid was built from its three directions. The clearest evidence of senior judgment.
Pair with · Sonnet
As the execution proof. Its working engine — parser, scoring, live multi-component — is the exhibit that shows the product's logic isn't hand-waved. It makes the case study credible.
Counterpoint · Opus
As the restraint and art-direction exhibit. The Doorway shows what "considered" looks like, and its guardrail chips show scope discipline made visible.
Don't crown one model — show the spread, then show your hybrid on top. Three strong-but-different outputs plus one human synthesis is a far better proof of "I can direct AI and out-judge it" than any single winning screen. If a reviewer forces a single pick: Fable, 9/10, with Sonnet's prototype attached as the execution receipt.
How to present this in a portfolio
This comparison is itself a case-study artifact. Framed well, it demonstrates the skill hiring managers most want to verify right now: not that you can generate with AI, but that you can direct and critique it.
- Frame it as "AI as a design material." The subject isn't Threadwise — it's your ability to use AI to design well. Say that in the first line.
- Lead with the scope-held finding. It answers the skeptic's first question ("doesn't AI just go off-brief?") with evidence before they ask it.
- Use the matrix as the scannable centrepiece. One table carries the whole comparison; it's the image a reviewer will screenshot.
- Show the drift analysis to prove critical reading. Anyone can prompt. The senior signal is knowing what to distrust — the five-dimension "why," the missing edge states.
- End on the hybrid as the payoff. The Doorway + Docket + Signal synthesis is your judgment layer — the human move the models set up but couldn't make.
- Keep the receipts. The nine artifacts plus this evaluation are the evidence base. Reference them; don't just assert the conclusions.
A suggested case-study spine
| Beat | What you show | What it proves |
|---|---|---|
| 1 | The frozen brief — scope, scoring, system, tone. | You can define and constrain the problem. |
| 2 | Three model runs, side by side. | You can use AI to explore breadth fast. |
| 3 | This evaluation and its matrix. | You can assess design work against criteria. |
| 4 | The drift you caught and why it mattered. | You read AI output critically, not credulously. |
| 5 | Your hybrid — Doorway structure, Docket rows, Signal rules. | You add the judgment AI can't. The payoff. |
Present it evidence-first and honest. The fact that you rate your own tools critically — scoring, flagging drift, naming what each still needs — and never claim "the AI did the design" is itself the most senior signal in the piece. Confidence with receipts beats a highlight reel.