# Qualitative public report writing guide

This guide covers the experimental person portrait that sits after the existing
taggers, summaries, score graders, and peak-moment extraction. It is modular and
is not part of the regular public-report generation pipeline yet.

## What the report is for

The report should let another person answer four questions quickly:

- How strong is this person, and in which ways?
- What do they repeatedly think about?
- What are they building or trying to change?
- Would I want to meet, hire, fund, work with, or learn from them?

Scores give the first read. Concrete evidence makes the person recognizable.
The report is not a personality horoscope, a ranking page with prose attached,
or a compressed biography.

## Page hierarchy

1. Put the person's name and one distinctive sentence first. Do not add a
   “Biggest spikes” card, overall level, composite human rank, or detached list
   of percentiles. Scores belong beside the metrics whose evidence explains
   them.
2. Follow with what the person is doing now and why it matters. Name the company,
   research direction, role, or exploration and explain it in one concrete
   sentence. Follow with the longer ambition only when the evidence supports it.
3. State what the person is looking for now. This may be collaborators, a role,
   a cofounder, a next direction, or people exploring the cutting edge of a
   field. Do not force everyone into a hiring or founder template.
4. State where they are based when a current public city or region is available.
   Keep it simple: city and state, region, or country. Do not infer a location
   from old schools, employers, or travel.
5. Follow with self-reported background and relevant prior work. Include current
   and recent roles from linked public sources, with dates and one concrete
   accomplishment. Credentials should not crowd out the work itself.
6. Do not add a separate scores-at-a-glance card or radar to the public portrait.
   The scored rows inside the body give each number its necessary context.
7. Organize the body around the existing important report sections, such as
   General Reasoning, Agency & Work, Growth & Information Intake, and
   Interpersonal.
8. Within each section, show the scored traits first. Judgment belongs under
   General Reasoning; it is not a separate top-level category.
9. After the traits, show the concrete portrait for that section: ideas,
   decisions, projects, reading, relationships, or other evidence that makes the
   scores intelligible.
10. Scores are compact orientation, not the main story. Do not add commentary to
   every number at the top.

The background should read as a compact trajectory. Use one bullet per major
period: duration, role or institution, and the concrete work that made the
period relevant. Include every college, but give degrees and work rather than
school names alone.

End with **Other relevant context**, reserved for the person's strongest signals:
rare ranks, competitive distinctions, publications, patents, major outcomes, or
other facts that sharply change how a reader understands their level. It is not
a miscellaneous bucket. Prefer three or four exceptional facts over a complete
list of memberships and activities.

## Top identity

The opening contains the person's public name and the most distinctive true
thing about their mission, trajectory, work, mind, or effect on other people.
When known, add their city or region as a quiet line. Do not put percentile
spikes in a separate opening card; the scored metric rows provide the necessary
context later.

The final sentence is not labelled “what is awesome about them.” Its content
should make that obvious. Prefer a duration, repeated choice, body of work, or
specific combination of ideas over praise.

Few-shot example:

> **Akshay Iyer**
>
> A solo founder who has spent ten months and four pivots pursuing one question:
> how to change what society measures so people keep developing and finding
> opportunity after AGI.

Bad:

> **Level 10 · Top 0.1% person**
>
> A visionary built to run something that would not exist without him.

The bad version collapses unlike traits into a judgment of the whole person and
replaces the concrete trajectory with promotion. A detached “Biggest spikes”
list has a milder version of the same problem: it repeats the scores before the
reader can see their evidence.

Every score shown anywhere in the report must map to a facet the pipeline
actually records. Never invent a plausible-sounding axis to complete a radar or
balance a layout. For example, use recorded OCEAN Agreeableness when relevant;
do not fabricate “Technical Affinity” from technical work or education.

## Choose the format that fits

Do not force every section into one visual template. Choose the smallest form
that fits the material:

- Use a heading and one or two short paragraphs for a central argument.
- Use a subheading and concise bullets when several distinct ideas need to be
  scanned.
- Use one compact evidence block when counts or a time window matter.
- Use short trait rows for score, percentile, claim, and expandable peak moments.

Avoid tables, two-column idea grids, divider lines between portrait items, giant
paragraphs, a stack of five micro-headings with one sentence each, and reports
made entirely of bullets. Whitespace and hierarchy should do most of the work.

Do not render substantive evidence in faint grey. Grey is for metadata such as
dates, provenance, or privacy notes—not claims, distributions, examples, or
topic counts. When a distribution contains several distinct facts, use a short
dark-text bullet list rather than one long secondary paragraph.

Do not surface aggregate grader histograms such as “70 of 77 conversations
scored 8/10 or above.” The facet score and percentile are the only aggregate
grading output a public reader needs. Use the space beneath them for behavioral
frequency, breadth, decisions, ideas, and concrete evidence. Individual
conversation grades may remain inside the private **See more** evidence because
there they annotate the exact moment being inspected.

## Do not re-explain the same fact

Give each important fact one canonical home. The opening owns the current work
and mission. Background owns the trajectory, dates, institutions, and major
credentials. Later sections may reference those facts only briefly and only to
show something specific to that section.

Bad Agency & Work copy:

> Over ten months he made four pivots toward human development after AGI. Today
> that mission is Polymath Society, which reads ChatGPT, Claude and coding
> histories. The long horizon includes BCI, governance and x-risk.

This repeats the opening, background and recurring-ideas sections without adding
evidence about agency.

Better:

- **Owns the whole loop:** writes the code, recruits candidates, runs work trials
  and gets employers to inspect the resulting reports.
- **Moves when the evidence changes:** each pivot removed a specific failed
  assumption, and the next product usually followed within days.
- **Creates signal before permission:** got five companies to review candidate
  reports before the current product had revenue.
- **Does not protect sunk work:** killed products with users or obvious financial
  upside when they failed the mission test.

When reusing evidence, compress it and change its function. “Four pivots in ten
months” is a trajectory in Background, evidence of fast updating in Agency, and
should not receive a third full explanation elsewhere.

## Evidence rules

Every public claim should be one of four things:

- **Supported:** directly present in the person's material.
- **Inference:** a careful synthesis from more than one piece of evidence.
- **Unsure:** plausible but not yet adequately investigated.
- **Rejected:** the person marked it wrong; do not quietly reintroduce it.

Preserve corrections and nuance from feedback notes. Investigate unsure items
with targeted retrieval. Do not turn a single conversation into a stable trait
or recurring interest.

When a claim is personal, show the trait but keep private receipts private.
Obfuscate non-public people's names. Public figures, works, companies, and
verifiable public facts can remain concrete.

## Recurring ideas versus one-off investigations

This distinction matters most.

An idea belongs under **Ideas he keeps coming back to** only when it appears
across multiple conversations or dates, changes a decision, shapes a project, or
is applied repeatedly. State the actual idea, not merely the topic.

For example:

- Post-AGI meaning: if intelligence and white-collar work become abundant, what
  will let people contribute, earn status, and feel needed?
- Measurement and human development: the SAT shaped schools, the bar exam shaped
  law school, and Leetcode reshaped computer-science education; changing the
  measure can change what people develop toward.
- Consciousness: whether sentience, world-modelling, and self-generated
  continuity are distinct properties.
- Meditation and non-attachment: using the Gita and Vedanta to pursue extreme
  ambition without making outcomes or destiny the basis of inner stability.

A one-off investigation belongs under **Other things he has explored**. Give it
one concrete line and do not inflate it into a conviction:

- Oligopolies in electric vehicles, commercial space, cloud computing, and
  automated manufacturing.
- Ultrasound-based non-invasive brain-computer interfaces.
- Longevity as a biological clock that might be reset.
- Whether universes exist partly to simulate their own possible futures.

## General thoughts

Use the public heading **General thoughts**. Do not call the section
“generalizations.” The useful idea from Harrison Qian's
[`generalizations`](https://harrisonqian.com/generalizations) and
[`notes`](https://harrisonqian.com/notes) pages is the content distinction, not
the label or layout:

- recurring ideas are durable questions or theories about the world;
- general thoughts are conclusions reached through reflection or
  experimentation and now used as working beliefs;
- other ideas and open questions are interesting but not yet established.

A general thought should sound like something the person genuinely learned and
would tell a friend. Use a short concrete label and one or two sentences. Do not
impose a small fixed cap: a sparse corpus may support three; a genuinely idea-rich
corpus may warrant ten to fifteen concise entries so the reader can see its real
range. Never pad a sparse record, and never hide a rich one merely to preserve a
symmetrical layout. Do not turn ideas into personality commentary, advice from
the report writer, or generic maxims that could describe anyone.

When additional strong ideas are omitted for space, say so only with evidence.
Prefer “The analyzed corpus contains 33 retained ideas and epiphanies; these are
the clearest public examples” to an unsupported “and many more.” Keep recurring,
endorsed ideas distinct from exploratory hypotheses and private reflections.

### Few-shot examples: general thoughts

Evidence: several work stalls ended as soon as a specific next action was named.

Bad:

> He is highly motivated but sometimes struggles with execution ambiguity.

Good:

> **Momentum:** work often stalls because the next action is ambiguous, not
> because motivation has disappeared. Make the next move concrete and motion
> returns.

Evidence: four product changes retained the same education and human-development
mission, with each change tied to new user or market evidence.

Bad:

> He is persistent and learns from pivots.

Good:

> **Pivots:** approaches can churn while the mission stays fixed. Four pivots
> toward the same problem can be convergence when each one removes a wrong
> assumption.

Evidence: work quality and emotional stability repeatedly improved with daily
company and worsened during isolated solo stretches.

Bad:

> He is socially regulated and benefits from community.

Good:

> **People and work:** daily intellectual company materially changes how well he
> functions. A people-rich environment works better than trying to become
> someone who needs nobody.

Evidence: the person repeatedly uses destiny language for motivation, catches it
distorting judgment, and relies on trusted people to challenge it.

Bad:

> He balances ambition with humility.

Good:

> **Ambition:** destiny-framing can generate enormous energy and still distort
> judgment. Preserve the ambition while keeping people and practices that catch
> the distortion.

Evidence: the person combined Greek love categories with attachment research to
separate projection from knowledge of a partner.

Bad:

> He thinks deeply about relationships.

Good:

> **Love and attachment:** eros can run on projection; philia requires accurate
> knowledge of the actual person. Attraction is not the same thing as evidence
> of compatibility.

Evidence: delayed outreach and decisions repeatedly centered on judgment,
rejection, or irreversibility rather than task difficulty.

Bad:

> He can overthink important decisions.

Good:

> **Irreversible asks:** procrastination often protects against judgment,
> rejection or a decision that cannot be taken back. More planning does not
> resolve that kind of fear; making the ask does.

### Few-shot examples: open thoughts

These remain one concise line because the evidence does not show a settled
belief:

- Oligopolies in electric vehicles, commercial space, cloud computing and
  automated manufacturing.
- Whether ultrasound can combine the speed missing from fNIRS with the spatial
  precision missing from EEG for non-invasive BCI.
- Whether ageing behaves like a biological clock that can be reset rather than
  an accumulation that can only be slowed.
- Whether universes might exist partly to model their own possible futures.

## Information intake

Information Intake is evidence about exposure, range, depth, and application.
Conversation count and message count do not discriminate between serious
reading and ordinary ChatGPT use.

Show counts first when the source supports them:

- books explicitly read or strongly recommended;
- named essays, blogs, and long-form series discussed;
- substantive ideas explored with AI in a stated indexed window;
- recurring works or authors returned to across separate conversations.

Counts must be reproducible, not estimates written for effect. If the evidence
shows at least 37, say 37+ rather than pretending the record is complete.
Separate what the person read from what an assistant merely recommended. State
the window for corpus-derived counts and the exclusion rule for logistics or
casual utility chats.

Never follow a book count with generic category prose. Name representative books
across fields so the reader can see the actual intellectual range. For example:

> At least thirty-seven books are directly named as read or strongly
> recommended. For judgment: *Poor Charlie's Almanack*, *The Black Swan*,
> *Antifragile*, and *Superforecasting*; economics and strategy: *The New China
> Playbook* and *The Art of Strategy*; AI and futurism: *Life 3.0* and *The
> Singularity Is Near*; founders and power: *The Mind of Napoleon*, *Titan*,
> *Elon Musk*, and *The Optimist*; philosophy: the *Bhagavad Gita*, *Ashtavakra
> Gita*, and *Meditations*.

The same concreteness applies to other sources:

- **Essays and blogs:** name Ben Kuhn, Sam Altman, Paul Graham, Hamming,
  LessWrong, or *Situational Awareness* when those are the actual sources.
- **Twitter:** say it is a morning feed and name the people and topics actually
  followed, such as Marc Andreessen, Sam Altman, new AI companies, funding,
  biotech, AI safety, and rationality.
- **Podcasts and YouTube:** name David Senra's *Founders* and the founders or
  historical figures repeatedly studied. If watch history is unavailable, do
  not invent a YouTube pattern.
- **Art:** name the artist and what the person takes from the work, such as Taylor
  Swift for concrete writing, Bleachers for love without sentimentality, The
  1975 for sound and angst, or John Mayer for guitar.
- **AI exploration:** state the number of substantive idea threads, the indexed
  time window, representative questions, and what was excluded.

Depth is shown by recurrence and use: returning to *The Mind of Napoleon* across
three conversations, revisiting *Life 3.0*, using the Gita in live questions
about work and attachment, or turning design reading into product standards.

### Information-intake calibration

Grade the person's information life, not whether they perform one culturally
preferred format. Asking questions about something they just read, arguing with
it through AI, connecting it to another source, or changing a decision all count
as processing.

- **10:** serious reading or learning several times a week or near-daily; many
  books, essays, blogs, papers, podcasts, feeds, or fields; active questioning,
  synthesis, or application. Many books and many blogs across fields is simply a
  10. Do not reserve 10 for literally reading every day.
- **9:** sustained phases of serious reading and questioning, with multi-week
  gaps or a narrower source range.
- **8:** occasional deliberate books, essays, blogs, or papers on top of active
  feeds, with some processing but inconsistent depth.
- **7:** mostly Twitter, YouTube recommendations, podcasts, or AI explanations,
  but real curiosity and learning are still visible.
- **5 or below:** sparse or passive exposure with little retained or pursued.

The public report should show the evidence for the score: lower-bound counts,
the time window, source mix, named examples across fields, recurrence, and any
visible application. The score alone is not an information footprint.

## General Reasoning

The canonical public rows are raw reasoning, judgment, bias resistance, and
taste. Prioritization can remain an internal grader but normally stays out of the
public portrait; its best evidence is usually already a judgment or agency
decision. Each public claim needs a concrete receipt or a defensible corpus
distribution.

Bad:

> He is strongest at finding hidden assumptions and challenging his own views.

Better:

> He forced a recurrent transformer through IIT's own rules until the theory had
> to call it conscious, then separated sentience from self-generated experience.

After the scored traits, add **Ideas they have been thinking about**. This is a
structured portrait, not another flat trait list and not a wall of bullets:

1. Give one or two load-bearing ideas a short developed paragraph each.
2. Put several durable but smaller questions under **Other recurring ideas**.
3. Put one-off or lightly explored questions under **Other things explored**.

Name the actual idea and its consequence. Keep recurring ideas separate from
smaller investigations. Do not use analyst commentary about what the person's
thinking style supposedly reveals when the underlying idea can be stated
directly.

For each scored row, show the *shape* of the quality, not a generic definition:

- **Raw reasoning:** breadth and frequency of substantive thought, followed by
  the person's most important or contrarian ideas in the structured block. Show
  both the dominant areas and the long tail of smaller investigations. For
  example, do not reduce a broad corpus to company-building, meditation and
  consciousness when it also contains market structure, BCI, neuroscience,
  longevity, AI safety, cosmology, design, art, history or relationship models.
  Use counts for the major areas when defensible; name the smaller fields without
  pretending each one recurs.
- **Judgment:** one or two high-consequence decisions, especially attractive
  options killed or directions changed quickly when evidence changed.
- **Bias resistance:** evidence that self-criticism and active search for
  disconfirming evidence happen repeatedly, plus one or two clear instances.
- **Taste:** where taste repeatedly appears—copy, positioning, demos, launch
  material, product decisions, visual design—and where the evidence is thin.
  Coding logs are valid evidence for taste when they contain the person's actual
  edits and reactions.

The Raw Reasoning row owns smaller, concrete reasoning threads that demonstrate
breadth but are not recurrent or consequential enough to become named ideas.
The Ideas section owns durable theses and repeated questions. Never duplicate a
small investigation in both places merely to make both sections look full.

Taste receipts must pass a zero-context test. Every example names:

1. the artifact being made and who it was for;
2. what the weaker version did;
3. the concrete line, structure, visual choice, or cut the person supplied; and
4. why that changed what the audience would understand or trust.

“Made claims scarcer to preserve credibility” fails because the reader does not
know which claims, on what page, or what changed. “The first public report had
eight top-percentile badges; he kept only the three rarest in the opening and
moved the rest beside their evidence” stands on its own. Never make a stranger
reconstruct the surrounding product conversation.

## Agency & Work

Write one mission, not an unordered list of interests.

Lead with the trajectory: over roughly ten months, four product pivots moved
toward the same mission. Name the products and what each taught. Then explain the
current implementation—Polymath Society—and how it expresses the larger thesis.

Future interests such as AI governance, non-invasive BCI, x-risk, or
philanthropic allocation are possible future arenas. Keep them subordinate to
the current work rather than presenting them as peer projects.

Avoid mechanical labels such as “path to the current idea” when the concrete
fact is stronger:

> In roughly ten months he moved from an AI tutor to Alphafeed, a mentored cohort,
> and AI-supervised work trials. Each product sharpened the same conclusion:
> whoever controls the outcome eventually shapes what people learn.

## Interpersonal

Keep rubric scores compact, including warmth, candor, conflict conduct, and
openness. Use “openness,” not “perspective-taking,” when that is the intended
trait.

Openness belongs here because willingness to consider conflicting views and
evidence is broadly important to relationships, collaboration, learning, and
decision quality—not merely because one person's corpus happened to contain a
large amount of evidence. When the evidence is strong, give it several concrete
instances rather than treating it as an optional novelty.

The portrait should make the interpersonal reality concrete: attachment,
loneliness, belonging, social regulation, the desire for daily intellectual
company, or the cofounder question. Explain the distinction between wanting an
executor and wanting someone who can share the intellectual and emotional
weight of the work.

Do not turn intimate material into spectacle. A public report can say that care
survived conflict while keeping identifying receipts private.

## Key motivation

The person's own account leads. Do not replace a stated motivation with an
analyst's synthesis of their behavior, mission, attachment patterns, or
personality. Behavioral evidence can add context only after the person's own
formulation is stated clearly.

For Akshay, the primary motivation is:

> The world has been exceptionally kind to him, and he wants to give back. His
> own fulfilment comes easily to him; the work is an attempt to help more people
> find purpose and opportunity.

Do not substitute “preserving ambition without destiny-framing,” “building a
people-rich life,” or “counterfactual impact” as the primary drive. Those may be
real tensions or decision rules, but they are not his stated reason for doing the
work.

## Voice

Write like a sharp personal website, not an evaluator narrating a subject.

- Use plain verbs and proper nouns.
- Lead with the concrete fact, count, title, decision, or repeated question.
- Prefer “Over ten months, four pivots moved toward one mission” to “His path to
  the current idea shows persistence.”
- Prefer named books to “works across judgment, psychology, and economics.”
- Prefer the actual IIT argument to “he finds hidden assumptions.”
- Keep claims short enough to scan, but do not split one coherent thought into
  many decorative fragments.
- Use third person in the outward-facing report. Preserve short verbatim quotes
  when they carry distinctive voice.

Ban empty praise, analyst self-commentary, and phrases that sound impressive but
do not add evidence: “decorative list,” “intellectual cluster,” “load-bearing”
without explaining what bears the load, “operating system,” “demonstrates
range,” “unusually strong,” or “exceptional” without the concrete act that earns
the word.

## What stays out

- Do not include a public failure-modes section in this portrait format.
- Do not include a “How to work with this person” section. It turns evidence into
  unsolicited instructions, repeats material already visible elsewhere, and
  makes the portrait sound managerial rather than human.
- Do not elevate a one-off chat into a recurring belief.
- Do not guess private community memberships, media habits, or reading.
- Do not present current and hypothetical future projects at the same hierarchy.
- Do not repeat the same evidence under several sections.
- Do not use generic summaries where names, counts, and moments are available.
- Do not show Prioritization or Implementation as default public rows. Preserve
  their useful evidence under Judgment, Agency, Self-diagnosis, or the relevant
  concrete portrait instead.

## How they work

This section is a compact deterministic summary, not a second qualitative
personality report. Show only measured work span, active days or days per active
week, focus/flow, AI parallelism, and related percentiles that the coding
analysis actually supports. Label the section **How they work (from coding
logs)**. Never name a specific coding tool or vendor in the public label.

For **Hours put in**, score the person's typical active work phase rather than
dividing active days by the entire archive span. Use days present in a typical
active week plus the usual first-to-last activity span:

- **10:** seven days at roughly 10+ hours, or six days at roughly 12+ hours;
- **9:** seven days at roughly 7+ hours, or six days at roughly 9–10 hours;
- **8:** six days at roughly 6–8 hours, or five days at roughly 10–12 hours;
- **7:** five days at roughly 8 hours, or four days at roughly 10 hours;
- **6:** five days at roughly 6 hours, four days at roughly 8 hours, or a
  comparable 30+ span-hours across the week.

Morning-to-night work counts as roughly a 12-hour daily span, but it does not
independently force a 10; days/week still matters.

Coding logs are the primary source and are a lower bound. When linked chat
summaries exist, check only conspicuous coding-gap days for clear GTM,
recruiting, meetings, research, or other work and backfill those days. Do not run
an exhaustive second analysis or infer work from an ambiguous conversation.

## Modular, token-efficient generation

Keep this portrait stage separate from the regular scoring pipeline while the
format is being tested. Reuse earlier work in this order:

1. score and percentile outputs;
2. facet summaries and peak moments;
3. conversation and source tags;
4. existing per-thread summaries;
5. targeted retrieval from raw chats only for counts, disputed claims, and
   concrete receipts.

Do not reread the entire corpus to write every section. First form candidate
claims from summaries and tags, then retrieve only the conversations needed to
verify recurrence, quote accurately, or establish a lower-bound count.

The module should preserve provenance for every rendered claim: source IDs,
dates, whether it is supported or inferred, recurrence count, privacy level, and
any user correction. The public renderer can stay concise because the evidence
remains available underneath.

### Extract while the initial tagger already has the conversation open

The initial broad tagger is also the per-conversation reader: it sees every
conversation once. During that same model call it emits a small
`portraitCandidates` sidecar alongside the title, evidence buckets, and
`consumed` sources. This is the normal production path. The deterministic
`portrait-evidence` stage merely joins those candidates into the shared
repository; it does not call a model.

Keep this sidecar raw and compact. It records a plain cluster, one concrete
observation, a few specifics, and an exact person-authored quote. When reading
or exploration produced a thought, preserve the full chain when present:
`inspiredBy` (source) → `observation` (what they concluded) → `application`
(what changed in their work, behavior, or worldview). A consumed source alone
belongs only in `consumed`; it is not a thought or epiphany.

The former rich selected-chat `portraitNotes` extractor remains available only
as an explicit diagnostic/backfill tool. It is not part of a normal run. This
avoids a second reading pass and its roughly $0.15-per-selected-conversation
cost. Do not put the full rich portrait schema back into the broad tagger: that
overloaded the cheap model in calibration. Recurrence, prominence, polished
writing, and section placement stay downstream.

The useful level is between a topic label and specialist shorthand:

- Too vague: “Interested in consciousness.”
- Too technical: “Separates sentience, world-modelling and self-generated
  continuity as orthogonal properties.”
- Right level: “Explored what makes conscious experience enriching, what might
  make an AI conscious, how meditation changes conscious experience, and
  theories such as integrated information theory.”

A single reader records evidence under a plain cluster. It never calls the
idea recurring or central. The deterministic collector counts distinct
conversations and dates before recurrence is available to the final writer.
Keep clusters broad and human (`Education`, `Consciousness`, `Meditation`,
`Company building`); the note title carries the narrower thought. Never use
empty clusters such as `intellectual`, `building`, or `systems thinking`.

General thoughts are a first-class evidence type. Preserve understandable
original ideas, distinctions, analogies, disagreements, changes of mind,
self-derived rules, and questions followed through their consequences. This is
how the report shows the texture and depth of the person's mind without generic
personality commentary.

Affinities are cross-cutting. Record what kinds of work, technical depth,
operating, institutional change, problems, people, communities, art, or ways of
living pull the person in, plus the concrete evidence for why. Do not force a
person into operator-versus-thinker binaries.

Current work carries a time status. A project active in an older conversation
is `active_at_time`, not automatically the person's work today. Preserve former
iterations as the path to the current mission and keep future directions below
the current implementation.

Every note must pass the stranger test: would an intelligent stranger learn a
specific, consequential fact about what this person thinks, wants, reads,
builds, values, enjoys, or does? Empty categories are correct. Incidental
mentions and assistant-only ideas are not portrait evidence. `other_notable`
is the escape hatch when something important falls outside the current schema.

### Public background in CLI onboarding

Ask once for LinkedIn, a personal website, and any additional public links
(comma-separated). Persist them locally before analysis and read the public
pages deterministically into the grounding digest. Do not wait until the share
screen to discover that the report lacks education and work context.

The CLI shows `Question X of Y` immediately above every input. It detects the
person's plan and corpus, states the resulting schedule, and runs it; never ask
the person to choose quota timing the program can decide itself.

## Canonical sample parity matrix

This table records the defects found by comparing a generated portrait with the
canonical sample and the owner's 2026-08-09 review. It is an acceptance
specification, not optional style advice.

| Surface | Failed draft | Canonical requirement | Acceptance check |
| --- | --- | --- | --- |
| Score authority | The portrait writer appeared to choose or reinterpret scores. | The final whole-person grader reads the evidence scoreboard and the full facet rubric. The portrait only renders those accepted scores. | Every trait score maps to a current `_person.json` facet. The writer cannot emit a free-standing score. |
| Final scoring model | The final disposition read used the same mini tier as per-chat grading and was split into facet shards. | Use the strongest subscription tier for the five whole-person judgments and grade all facets in a dimension together. | Provenance says Claude Opus or Codex `gpt-5.5`, with `gradingMode: full-dimension`. |
| Background structure | Jobs, universities, and projects were flattened into one list. | Work experience, education, ventures, and other distinctions are separate kinds. Jobs name role, company, and dates. Education names institution, degree, and dates. | At least one typed work entry and one typed education entry when both exist in supplied sources. |
| Current work | The draft said it was an npm local tool but did not explain the product or consequence. | Explain the problem, mechanism, immediate user benefit, downstream opportunity, and how the work could change education or labor markets. Packaging and privacy are supporting facts, not the thesis. | A stranger can explain what it does, for whom, why it matters now, and what it could unlock. |
| Trait rows | A score appeared with a generic one-sentence summary. | Put the accepted score and percentile beside the trait name, then state the precise shape of the capability or disposition. | No duplicate `Score X` sentence. No trait body that could fit any competent person. |
| Proof bullets | Bullets were chart titles plus dates. | Each bullet names the situation, the person's concrete move, the alternative or constraint, and the outcome or consequence. | Reject title-only, date-only, and noun-phrase bullets. Each receipt must stand alone for a new reader. |
| Raw reasoning | Two examples stood in for the whole ability read. | Show breadth with lower-bound counts, recurring domains, smaller fields, and several peak moves. | At least three stranger-readable breadth points plus concrete peak receipts. |
| Judgment | The explanation stayed inside one topic and did not show the many-decision pattern. | Synthesize several high-stakes decisions across work and personal contexts, including costly calls, contrary evidence, and recurring lapses. | The body explains why the full pattern earns the rung, not why one highlight is impressive. |
| Bias resistance | Asking for criticism was treated as enough. | Show self-generated disconfirmation, belief movement, and decisions changed by counterevidence, while naming identity-loaded exceptions. | At least two complete proof receipts and an honest description of the limiting pattern. |
| Taste | Shorthand such as a product or copy title replaced the actual aesthetic choice. | Name the artifact and audience, weaker version, exact direction chosen, and why it landed better. | Every Taste receipt passes the stranger test. |
| Ideas | A short recurring list omitted the broader intellectual record. | Separate recurring ideas from one-off thoughts and hypotheses. Preserve a grounded retained-idea count. | Rich corpora normally show ten or more diverse ideas, with recurrence labeled honestly. |
| Information footprint | The Information intake score had no visible substrate. | Show lower-bound counts for books, essays or blogs, fields, and other sources, followed by named examples grouped by type. | A structured footprint includes at least three numeric lower bounds and named source groups. |
| Interpersonal evidence | Privacy handling removed all concrete proof. | Keep privacy-safe actions and decisions while removing names, quotes, and identifying details of other people. | Warmth, candor, conflict conduct, and openness each retain concrete receipts without exposing private people. |
| How They Work | The section read a stale five-day aggregate and omitted score badges. | Reuse the coding profile's own calibrated score, percentile, and full-history deterministic metric for push, hours, parallelism, and focus. | Four rows, each with score, percentile, exact number, window, and a corpus-freshness gate. |
| Sparse windows | A partial recent window silently replaced the full record. | Coding logs are a lower bound, but a stale or materially thinner aggregate cannot be presented as the full history. | Aggregate generation must not predate its session, flow, or concurrency inputs; active-day counts must agree with the compiled profile. |
| Acceptance | Structural presence alone counted as success. | Compare section by section with the canonical sample for depth, self-containment, proof density, and numerical grounding. | Generation fails before publish when proof, idea, footprint, background-kind, or work-metric gates are missed. |

## Final quality check

Before showing a draft, verify:

- Can someone understand the person's level from the first screen?
- Does each trait have a concrete moment, not generic praise?
- Are recurring ideas genuinely recurrent and stated as ideas rather than topic
  labels?
- Are one-off investigations visibly smaller?
- Does Information Intake show counts, time windows, named sources, and examples
  across fields?
- Are assistant recommendations excluded from reading counts?
- Is the current mission clearly above possible future arenas?
- Are wrong claims removed and unsure claims investigated?
- Are paragraphs short, bullets used only when useful, and two-column idea grids
  absent?
- Could another thoughtful person decide whether they want to meet or work with
  this person after reading it?
