Skip to content
Threadwise.
Home Comparison

AI-Assisted Product Design · Case Study CompanionThreadwise v0.1

One frozen brief,
three model minds.

The same Threadwise foundation went to three Claude Design models — Sonnet, Opus, and Fable — each running the same staged process: propose three directions, resolve one, build hi-fi screens, then a working prototype. This is the evaluation: what held, what drifted, and where a designer's judgment still had to close the gap.

3 models compared 4 build stages each 12 evaluation criteria 9 artifacts + source code reviewed
Sonnet
The Engineer
completeness & working systems
8.5/ 10

Most complete and the only prototype with a live scoring engine — including the multi-component case none of the others wired.

Opus
The Minimalist
restraint & visual discipline
8.0/ 10

The cleanest, most consistent execution and the tightest scope guardrails — but the thinnest coverage of edge cases.

Fable
The Design Lead
judgment & synthesis
9.0/ 10

The strongest design thinking and the most principled honesty — and the three directions the final hybrid was actually built from.

The result

None of the three broke scope. No model added comparison, accounts, a wardrobe, streaks, or "buy this instead." Every difference below is about craft, completeness, and judgment — not hallucination. Against a well-frozen brief, capable models behave. The designer's real work moves up the stack.

01

Executive summary

Three models received identical Threadwise foundation documents and a near-identical staged prompt sequence. The experiment tests a simple, useful question for anyone using AI in product design: given a tight, well-specified brief, do capable models stay inside it — and where do they still need a human?

The reassuring headline is that all three respected the frozen MVP. Each kept Threadwise a single-garment tool, each explicitly excluded the Phase 2 comparison flow, and none invented features, accounts, or purchase suggestions. Two of the three even restated the guardrails on-screen ("single garment," "Phase 2 excluded," "no suggestions") before designing a pixel.

The differences showed up one layer down. Sonnet optimized for completeness and working systems: 18 fully specified screens and the only prototype with a real parser and scoring engine, including live multi-component handling. Opus optimized for restraint: the calmest, most internally consistent visual system and the strictest scope discipline, at the cost of several edge states it simply didn't draw. Fable optimized for design judgment: the richest exploration, an explicit "held constant" frame, per-direction interaction specs, and the most principled honesty behaviour — it refuses to invent a score when a label is only partly readable.

Key conclusion

Model choice here is a choice of strength, not correctness. Fable is the strongest single body of design thinking (and the directions the final hybrid was built on); Sonnet is the strongest proof of functional execution; Opus is the strongest demonstration of restraint. The most honest recommendation is to use all three as evidence of different senior skills — with the human synthesis on top as the real payoff.

02

Experiment setup

A controlled comparison: same inputs, same process, three model "brains." The only variable that changed was the model doing the work.

The four staged deliverables, identical across models
StageAskWhat it tests
01  Three directionsPropose three distinct hi-fi visual directions for the frozen MVP.Range of design thinking; can it stay in scope while exploring?
02  Resolve oneSelect and commit to a single direction (or a deliberate hybrid).Editorial judgment; can it choose and justify?
03  Hi-fi screensBuild every MVP screen and state in the chosen direction.Completeness; edge-case coverage; design-system fidelity.
04  PrototypeWire an interactive, tap-through version of the flow.Interaction logic; does the pattern actually work when clicked?
Assumption

The brief states the prompts were "nearly the same" across models. The prompts themselves were not part of this review — only the nine output artifacts and their underlying code. Where a model's stage intent is inferred, it is drawn directly from that artifact's own labelling (each model explicitly named its stage and directions). Scores reflect the evidence in the artifacts, spot-checked against the source documents; they are a designer's calibrated judgment, not an absolute metric.

03

What the foundation controlled

The Threadwise pack froze six things before any model touched it. These are the yardsticks every output is measured against — and the source of the one genuine ambiguity the models had to navigate.

The frozen inputs and what each governed
ControlWhat it locked
MVP scopeText-first, single garment, sustainability-led, with a durability-and-care layer. Comparison of 2–5 garments is Phase 2, explicitly out of the MVP. No accounts, wardrobe, brand ratings, notifications, or feeds.
Scoring logicA 0–100 Fabric Composition Score from a fixed fiber base-score table (linen 90 → elastane 20), weighted by percentage, with two small blend penalties (−5 mixed, −5 complex). An "illustrative, meant to guide" cue and a permanent "what this can't see" line.
Design systemThe "festival, not a sacrifice" heritage system: the jharokha arch motif, Heritage Pink #8B2149, cream #FBF4E6, Rozha One display serif over Hind body over IBM Plex Mono. "Green is earned," one accent per screen, never colour alone.
Hi-fi behaviourThe result is a card that rises over a dimmed Home. An eight-zone anatomy: verdict + score, honesty cue, one reason line (peek ends here), breakdown, end-of-life/microplastic, durability/care, an inline "why this score," and one calm action.
Interaction rulesPeek → expand → "why," all tap-driven, no drag physics. One overlay card maximum. "Why" is an inline accordion, never a second sheet. About, Error, and Compare are full pages — never dressed as result cards. Degrades to a single panel on desktop.
Content toneCalm, plain, warm, non-judgmental. "Better to skip," never "avoid" or "bad." "Breaks down," not "biodegrades." State the verdict and step back — no upsell, no guilt.
One ambiguity the brief left open

The foundation is not perfectly self-consistent, and this matters for judging the models fairly. The Source of Truth and Hi-Fi Spec describe three calm bands (Good to buy / Okay / Better to skip) and a "why" that opens five lifecycle dimensions. The dedicated Fabric Composition Score document narrows the MVP to composition only, uses five bands, a three-line explanation, and parks the lifecycle dimensions for v2. No model created this tension — but every model had to resolve it, and they resolved it differently. Watch it recur in the drift analysis.

04

Model-by-model analysis

Each model, read across all four stages and its underlying code. Same rubric, three temperaments: the engineer, the minimalist, and the design lead.

Design Model

Sonnet · the engineer

8.5 / 10

Sonnet treated the brief as a system to be fully built. Its three directions vary by density and craft — Minimal & fast, Score-card focused, Guided explanation — and it resolved them into a single "Plan A" hybrid, then shipped the most complete artifact set of the three.

01 Directions
Three directions on a density/teaching axis, 8 states each, explicitly "no comparison, no suggestions."
03 Hi-fi
18 screens — every edge case drawn: invalid input, partial read, multi-component, unrecognized, scan-fail, permission-denied.
04 Prototype
Live parser + weighted-average engine using the exact fiber table; the only prototype that computes multi-component verdicts.

Where it followed the source well

  • Embedded the frozen fiber base-score table verbatim (linen 90 … elastane 20) with synonym handling (rayon, polyamide, lycra).
  • Handled all six named edge cases, including the multi-component "verdict reflects the shell, lining shown separately" rule.
  • Held the card pattern precisely: tap-driven peek → expand → inline "why," About and errors as full pages.
  • Tone stayed calm and factual — the Skip variant reads "…working against you on fabric alone," never "bad."

Where it drifted

  • Used the five-band scale in its directions but the three-band scale in its hi-fi — an internal naming inconsistency to reconcile.
  • Showed the five lifecycle "why" bars, which the composition-only scoring doc parks for v2 — mitigated by a "not five separate scores" caption.
  • By its own admission the plainest direction "could be any utility app" — the least distinctive on brand of the three.
▲ Strongest contribution

Functional completeness. It is the only submission where you can type "70% wool, 30% polyester / lining 100% polyester" and get a correct, live, multi-component verdict — the clearest proof the product's logic actually works.

▼ Weakest area

Visual distinctiveness. The resolved direction is clean and legible but the least characterful; it spends its effort on coverage and correctness rather than making Threadwise feel unmistakable.

Human correction required

Pick one band scheme and propagate it; decide whether the "why" should expose five dimensions or hold to composition-only; push one notch further on brand expression if this direction is taken.

Portfolio-worthy ● Yes Strongest single artifact for proving the system works end to end.
Design Model

Opus · the minimalist

8.0 / 10

Opus treated the brief as something to execute with restraint. Its three directions — The Doorway, The Care Label, Calm Focus — are the most confidently art-directed, and it carried the frozen guardrails onto the canvas as literal chips before designing a thing.

01 Directions
Three art-directed directions with explicit "held constant" chips: single garment · card over Home · tap-driven · no suggestions · Phase 2 excluded.
03 Hi-fi
Resolved on "The Doorway." The cleanest happy path + three verdict variants (linen 90, cotton-poly 62, polyester 35) + recovery + About.
04 Prototype
A well-structured engine (regex fiber matcher + separate score table, strong synonyms); happy path + variants + recovery wired cleanly.

Where it followed the source well

  • Most internally consistent band model — three calm bands with the Source-of-Truth ranges (70–100 / 45–69 / 0–44), held across every stage.
  • Used the trace-synthetic honesty line almost verbatim ("…can't be recycled as a natural fiber") and the labelled fact rows (END OF LIFE / MICROPLASTIC / ON SKIN / CHECK YOURSELF).
  • Named the card-pattern boundary out loud: "Error and About are full screens, never result cards."
  • Highest design-token fidelity — leaned fully into the arch and heritage system without inventing off-brand colour.

Where it drifted

  • Drift by omission: several edge states the spec calls its "strongest-specified area" — multi-component, partial read, invalid input, permission-denied, low-light — are not drawn or wired.
  • Also showed the five "why" bars — the same tension with the composition-only scoring doc.
  • The Doorway's ceremony carries a cost its own notes flag: ~90px of arch header before any facts, which strains on rapid rack re-checks.
▲ Strongest contribution

Visual discipline and consistency. The Doorway is the most polished single direction, and the band-and-token system is the most coherent of the three — the best evidence of taste and restraint.

▼ Weakest area

Completeness. It shows the product at its best but not at its messiest; a designer adopting it inherits the job of drawing every unhappy path the others already covered.

Human correction required

Backfill the missing edge states (multi-component, partial, permission, low-light); resolve the "why" content against the frozen scoring model; pressure-test the arch header for the fast aisle case.

Portfolio-worthy ● Yes Strongest single artifact for proving restraint and art direction.
Design Model

Fable · the design lead

9.0 / 10

Fable treated the brief as a design problem to reason about out loud. It shipped the richest exploration — a shared "frame every direction walks," per-direction interaction specs, and a closing synthesis that reads each direction as a strategic bet.

01 Directions
Three directions — The Doorway, The Docket, The Signal — each with concept, layout, hierarchy, interaction notes (timing, scrim, reduced-motion, targets), strengths, risks, in-store read.
03 Hi-fi
Resolved on "The Signal." 19 screens including a keyboard state and fiber-aware suggestions; the most careful edge handling of the three.
04 Prototype
Functional engine that uniquely wires the principled partial-read refusal live; full tap-driven card lifecycle with scrim dismiss.

Where it followed the source well

  • Most faithful to the scoring doc's language — used its exact recommendation labels ("solid pick," "weak fabric choice," "probably skip").
  • Best embodiment of the honesty principle: on a partly-unreadable label it shows "Can't fully score … the rest is a guess we won't make. No score is better than a made-up one." — it refuses to invent a number.
  • Locked the frozen words explicitly ("fiber" not "fibre," calm verdicts never "avoid," cue rides with the score, no scoring math in the flow) as a stated constraint.
  • Anticipated the exact hybridization path — "the Signal's peek with the Docket's expanded rows" — before a direction was even chosen.

Where it drifted

  • Declared "no scoring math in the user-facing flow" in its directions, then reintroduced the five "why" bars in its hi-fi — the same latent tension, and a small inconsistency with its own rule.
  • Blends three-band headline words with five-band sublabels, so some screens carry two verdict labels at once — thoughtful but slightly redundant.
  • Its prototype wires partial and unrecognized states but not the multi-component case Sonnet handles live.
▲ Strongest contribution

Design judgment and narrative. It doesn't just produce screens, it produces an argument — and its honesty behaviour (refusing a made-up score) is the single most on-brand decision in the whole experiment.

▼ Weakest area

Discipline in its own detail. The double-labelling and the directions-vs-hi-fi "why" reversal show it occasionally out-reasoning its own constraints; it needs a light editorial pass to prune.

Human correction required

Resolve the "why" content once (its directions had the more principled instinct); prune to a single verdict label per screen; add the multi-component case to match its own thoroughness elsewhere.

Portfolio-worthy ● Yes — lead with it Strongest artifact for senior design thinking, and the basis of the final hybrid.
05

Comparison matrix

All twelve criteria at a glance. The bar shows strength; the colour shows which model. Read a row to compare the field on one dimension, or a column to read one model's shape.

Strength: Strong Solid Partial Watch
Sonnet · Opus · Fable across 12 criteria
Criterion Sonnet Opus Fable
MVP scope control01
Strong
Single garment held; "no comparison, no suggestions" stated up front.
Strong
"Phase 2 excluded" carried onto the canvas as a literal guardrail chip.
Strong
Froze scope + words in a dedicated "held constant" frame.
Product clarity02
Strong
"A huge score, one sentence, done" — unmissable value.
Strong
"Answer in one breath, then depth on tap" reads instantly.
Strong
Poster-scale question framing; the answer legible before a word is read.
In-store usability03
Strong
Fastest read at every step; quiet always-available exit.
Solid
Beautiful, but the arch ceremony adds overhead on rapid re-checks.
Strong
The Signal is built for the two-second, one-handed aisle glance.
Result-card fidelity04
Strong
Peek → expand → inline "why," one overlay, tap-driven throughout.
Strong
Named the boundary: About/Error as full pages, never cards.
Strong
Pinned verdict head, inline "why," scrim-to-dismiss — spec-exact.
Visual design quality05
Solid
Clean and legible, but self-described as the least "branded."
Strong
The most polished, considered surfaces of the three.
Strong
Three fully-realised design arguments, each unmistakable.
Design-system alignment06
Strong
Correct tokens + font stack; no off-brand colour.
Strong
Highest token fidelity; fully committed to the arch system.
Strong
Even matched the documentation pack's own chrome.
Content tone07
Strong
Calm and factual; Skip never reads as blame.
Strong
Honesty cue + "check yourself" line held consistently.
Strong
Warmest, most human copy; "no score is better than a made-up one."
Interaction quality08
Strong
Wired the most states, including live edge cases + dev triggers.
Solid
Cleanest happy-path wiring; fewer states reachable.
Strong
Best-specified motion (timing, scrim, reduced-motion) and wired refusal.
Edge-case coverage09
Strong
All six edges drawn and wired, incl. live multi-component.
Partial
Missing multi-component, partial, invalid, permission, low-light.
Strong
All edges + extras; most principled partial-read handling.
AI drift resistance10 · higher = less drift
Solid
No invention; band-scheme inconsistency + five-bar "why."
Strong
Most consistent; drift is by omission, not invention.
Solid
No invention; out-reasoned its own "no math" rule once.
Low human correction11 · higher = less needed
Solid
Reconcile bands; decide "why" content; lift brand a notch.
Partial
Inherits the most backfill — every missing edge state.
Solid
Prune double-labels; settle "why"; add multi-component.
Portfolio strength12
Strong
Proof the system works — the execution exhibit.
Solid
Proof of restraint — the art-direction exhibit.
Strong
Proof of judgment — the design-thinking exhibit, and the hybrid's source.
Overall score 8.5 / 10 8.0 / 10 9.0 / 10
↑ Back to top
06

Key findings

What this experiment revealed about using AI as a design material — beyond the three individual scorecards.

Finding 01 · The big one

Scope held everywhere. Against a tightly frozen brief, all three models stayed inside it — no comparison flow, no accounts, no wardrobe, no gamification, no upsell. The classic fear ("the AI will run off and redesign my product") did not materialise once. A well-written brief is the most effective guardrail there is.

02 · Capability = completeness + judgment

The gap between models was never the happy path — all three nailed that. It was coverage of the hard 20% (edge cases) and the quality of reasoning about tradeoffs. That's where "strong" pulled away from "adequate."

03 · The brief's own gaps propagate

Where the foundation contradicted itself (three bands vs five; composition-only vs five dimensions), each model faithfully followed one source doc. Faithfulness to a conflicting spec looks like drift but isn't — it's a spec problem only a human can settle.

04 · Working prototypes are table stakes

None of the three shipped a clickable mockup. All built real parsers and scoring engines from the written fiber table — Sonnet even computes multi-component verdicts live. The bar for "prototype" has moved from "looks interactive" to "is interactive."

05 · Edge cases are the tell

Same peek-expand-why on the happy path; wildly different unhappy paths. If you want to tell a strong model output from a passable one quickly, look at what it does when the label is torn, blurry, unnamed, or lined.

Finding 06 · Values are harder to prompt than pixels

Fable's refusal to invent a score when a label is only partly readable — "no score is better than a made-up one" — is a values decision, not a UI decision. It's the hardest thing to specify in a prompt and the most worth noticing when a model does it unasked. That instinct, more than any layout, is what "on-brand" actually means for an honest-broker product.

07

AI drift patterns

Where the models overreached, over-explained, or drifted visually. The honest read: drift here was mild, and mostly not hallucination — it was faithfulness to a brief that quietly disagreed with itself.

Observed drift, by pattern
PatternWhere it showed upSeverity
Five-dimension "why" All three, in hi-fi. The "why this score" opens five lifecycle bars (Climate/Water/Land/End of life/On skin) — faithful to the Hi-Fi Spec, but the composition-only scoring doc parks these for v2. All three softened it with a "not five separate scores" caption.
Low
Band-scheme divergence Sonnet used five bands then three; Fable blended three headline words with five sublabels; Opus stayed three throughout. No model invented bands — they chose among the brief's own two schemes.
Low
Out-reasoning its own rule Fable declared "no scoring math in the user-facing flow," then reintroduced the five "why" bars in hi-fi. A small self-inconsistency from a model thinking harder than its own constraint.
Low
Verbosity / double labels Fable carries two verdict labels on some screens (e.g. "PROBABLY SKIP" + "weak fabric choice"). Reads as thoroughness, but adds words a glance-first product doesn't need.
Low
Drift by omission Opus left several spec-emphasised edge states undrawn (multi-component, partial, invalid input, permission, low-light). The most consequential drift here — invisible until you audit for what's absent rather than what's wrong.
Medium
Just as important · what did NOT happen

No invented features. No sneaking in the comparison flow. No "buy this instead." No raw scoring math dumped on the user. No off-brand colour. The failure modes people most associate with AI-generated design were absent across all three — the drift that remained was subtle and interpretive, not reckless.

08

Where human judgment was necessary

The models widened the option space and built fast. Narrowing it, reconciling the brief's contradictions, and making the taste-and-values calls stayed human work. This is the real division of labour the case study demonstrates.

The pattern worth naming

AI moved the designer's work up the stack: less pixel-pushing, more direction, critique, reconciliation, and judgment. The skill on display in this case study is not "made screens with AI" — it's directed three models, read their output critically, and added the judgment they couldn't.

09

Recommendation

If the portfolio needs a single "strongest output," it's Fable — narrowly. But the more honest and more impressive move is to use the three together, each as evidence of a different senior competency.

Lead with · Fable

For depth of design thinking, the most principled honesty behaviour, and the fact that the final hybrid was built from its three directions. The clearest evidence of senior judgment.

Pair with · Sonnet

As the execution proof. Its working engine — parser, scoring, live multi-component — is the exhibit that shows the product's logic isn't hand-waved. It makes the case study credible.

Counterpoint · Opus

As the restraint and art-direction exhibit. The Doorway shows what "considered" looks like, and its guardrail chips show scope discipline made visible.

The recommendation, in one line

Don't crown one model — show the spread, then show your hybrid on top. Three strong-but-different outputs plus one human synthesis is a far better proof of "I can direct AI and out-judge it" than any single winning screen. If a reviewer forces a single pick: Fable, 9/10, with Sonnet's prototype attached as the execution receipt.

10

How to present this in a portfolio

This comparison is itself a case-study artifact. Framed well, it demonstrates the skill hiring managers most want to verify right now: not that you can generate with AI, but that you can direct and critique it.

A suggested case-study spine

BeatWhat you showWhat it proves
1The frozen brief — scope, scoring, system, tone.You can define and constrain the problem.
2Three model runs, side by side.You can use AI to explore breadth fast.
3This evaluation and its matrix.You can assess design work against criteria.
4The drift you caught and why it mattered.You read AI output critically, not credulously.
5Your hybrid — Doorway structure, Docket rows, Signal rules.You add the judgment AI can't. The payoff.
On tone

Present it evidence-first and honest. The fact that you rate your own tools critically — scoring, flagging drift, naming what each still needs — and never claim "the AI did the design" is itself the most senior signal in the piece. Confidence with receipts beats a highlight reel.

↑ Back to top