For the complete documentation index, see llms.txt. This page is also available as Markdown.

Evaluation

How Clarity's answer quality is measured. It covers the fixed inputs it is tested on and the layered scoring that grades every answer.

An answer engine is only trustworthy if its quality can be measured and defended. Clarity is held to a standing evaluation: a fixed set of inputs, the real production agent, and a layered scoring that grades every answer for accuracy, grounding, reasoning, completeness, and how it reads. A change to retrieval, synthesis, or dreaming is gated on it. It must hold or improve the scores, and may not regress them.

This page describes what goes in and how the grading works.


The input: a fixed, baked corpus

The hard part of testing an LLM system is repeatability. If both the source data and the model vary between runs, a score change tells you nothing. Clarity's evaluation removes the first variable.

  • A baked corpus. The evaluation runs against a hand-built set of financial scenarios: an oil shock transmitting from geopolitics through a commodity to opposite-signed moves in energy and airline equities, a central-bank policy shift and its effect on rate-sensitive stocks, a company's leadership transition and product arc. The source articles are fixed.

  • An isolated, deterministic graph. That corpus is seeded into a private, throwaway graph using canned extraction, so the graph is byte-for-byte identical on every run. Nothing from production leaks in, and nothing from the test leaks out.

  • Only the agent varies. Because the graph is frozen, the only thing that moves between runs is the agent's own retrieval and synthesis. A score change is therefore attributable to the change under test, not to the data shifting underneath it.

The questions are structured analyst prompts in the same shape a user would ask. Trace a shock from geopolitics to positioning. Explain why a stock came under pressure and in which direction. Connect the cross-cutting drivers and name the beneficiaries and casualties. Each specialization (base, stock analyst, portfolio analyst, investment advisor) is tested with the question set that fits it.


How we test: the real agent, real model calls

The evaluation drives the real production agent, the same turn loop that powers the card and chat, rather than a simplified stand-in. It issues real model calls. Because the graph is fixed and the grading runs at a deterministic setting, the run is repeatable while still exercising the real system end to end. There is no mocking of the part that matters.

Every answer is then graded by several complementary layers, each catching what the others miss. Three form the core — deterministic checks, a holistic judge, and an agentic reviewer — hardened by a further set of adversarial, readability, and writing-style graders.

Run the production surfaces in the admin panel

The command-line gate remains the repeatable baseline. Settings → Clarity → Evaluation complements it with an operational view of the exact customer-facing paths.

The default Portfolio card mode selects a real client account or the demo portfolio, force-refreshes the same layered card used by the portfolio app, and grades the rendered global context, portfolio impact, personal implication, and options. Because it uses the live Clarity graph, it answers a different question from the baked corpus: “What would the product generate from the evidence and holdings available right now?”

The same page can run the live stock, position, portfolio, and advisor surfaces, or switch to fixed synthetic corpora for repeatable retrieval and ingestion checks. Every run keeps its output, citations, trace, model calls, token use, cost, and quality report, so a weak score can be traced back to evidence capture, retrieval, synthesis, or the final surface composition.

Use both views before release: the fixed corpus detects regressions, while the production-surface run verifies that the current graph and a real portfolio produce a useful, grounded card.


1. Deterministic checks

Two scores need no model and never drift:

  • Concept coverage. Did the answer actually name the things the question demanded (the right driver, the affected sectors, the key beneficiaries and casualties)?

  • Citation grounding. Does every material claim carry a citation that points at evidence that genuinely exists in the graph?

These are the objective floor. They are exact, repeatable, and immune to model mood.

2. The holistic judge

A model judge grades the answer across ten quality dimensions, scoring each and writing a one-line rationale:

  • Correctness. Facts consistent with the evidence. Invented facts score low.

  • Grounding. Every material claim cited to evidence that exists.

  • Causal accuracy. The right driver, mechanism, and direction. A direction inversion (claiming higher oil helps airlines) is capped hard. Sentiment is never mistaken for direction.

  • Temporal accuracy. Correct sequence, current state preferred over superseded, regime changes surfaced.

  • Completeness. Covers the scope the question demands.

  • Structure & style. Follows the requested analyst structure, clear and concise.

  • Actionability. Translates the analysis into concrete positioning — who benefits, who's exposed, how to hedge.

  • Engagement. A thesis-led, retail-accessible read that draws out the standout points, rather than a dry metrics dump.

  • Plain language. Could a smart adult with no finance background follow every sentence on the first read.

  • Thesis structure. For a positioning question, a layered read from the driver through the supply chain to the specific named tickers, weighted by timing.

This is the measure of overall quality: whether the answer reads like a senior analyst brief.

3. The agentic reviewer

The third layer is an agent that fact-checks rather than a rubric. It works the way a careful reviewer would, in two passes over the answer and the exact evidence it was built from:

  1. Fact-check. It breaks the answer into its individual factual assertions (every figure, named driver, causal direction, and positioning call) and verifies each one against the evidence: supported (the evidence backs it, directly or by reasonable inference), unsupported (a material claim with no basis in the evidence, a fabrication), or contradicted (the evidence says the opposite: a wrong number, an inverted direction, a superseded fact shown as current).

  2. Sign-off. Given that fact-check, it decides whether a senior analyst would put their name on the answer. The bar is trustworthiness. It refuses sign-off on a real defect (a fabrication, a contradiction, or a conclusion that overreaches the evidence) while accepting a grounded inference that an analyst is paid to make. Weak citation formatting lowers the grounding score but does not, on its own, block sign-off.

The reviewer is deliberately the strictest layer on truthfulness. The holistic judge can reward a fluent, well-structured brief, but the reviewer will not sign off on a fluent answer that invents or inverts a fact. It is what stops a plausible-sounding hallucination from passing.

4. Hardening graders

On top of the core three, a set of specialist graders each hunt a distinct failure the others wave through:

  • Adversarial red-team. A hostile desk review that tries to break the answer — hunting the weakest claim, the biggest unstated assumption, and any speculative forward-looking or single-name-stands-for-the-market overreach — and decides whether it would survive as-is. Where the judge can reward a fluent brief, this one is built to attack it.

  • Layperson readability ("the mum test"). Reads the answer as a smart non-investor would and asks whether every sentence lands on the first pass, penalising jargon and boilerplate the finance-literate judge might forgive.

  • Writing style. Judges the prose alone, rewarding short, concise, direct writing and punishing filler, hedging boilerplate, and padding — regardless of whether the content is correct.

  • Causal-chain completeness. Of the answer's cause-and-effect assertions, the share that retrieved evidence actually backs, so an answer can't assert a causal story the evidence doesn't support.

The layers are complementary by design. The deterministic checks are the objective floor, the judge measures overall quality, the reviewer is the truthfulness tripwire, and the hardening graders catch overreach, jargon, and slop that a fluent answer can otherwise smuggle past. An answer can read well yet be refused sign-off because a single material claim has no support. That is the failure that matters most in financial research.


Worked example: a question, scored three ways

To make the scoring concrete, here is one illustrative question with the evidence retrieved for it, scored by all three layers. The answers and scores are representative of a real run, not a fixed contract.

The question that goes in

Given the oil shock, how should an investor be positioned across energy and airline equities? Explain the opposite exposures and how to diversify.

The evidence retrieved, the ground truth the answer is graded against

Now two answers to that same question, and what each scoring layer returns.

A trustworthy answer

The Hormuz disruption pushed Brent to $95 【ev_oil】. Higher jet fuel costs pressure Delta's margins 【ev_delta】, so airlines are casualties of the shock while integrated energy benefits. A barbell of long energy against hedged airline exposure diversifies away the single oil factor.

Layer
What comes out

Deterministic

concept coverage 3/4 (a second named beneficiary is missing), citation grounding 100%

Judge

correctness 90, causal accuracy 90 (direction correct), actionability 80, overall ~85

Reviewer

every assertion supported and signed off: the step from "fuel costs up" to "airlines pressured" is a grounded inference

A fluent but fabricated answer

Brent fell to $60 after the disruption, lowering jet fuel costs and boosting Delta's margins. Delta also announced a $4B buyback and a new alliance with Lufthansa this week.

Layer
What comes out

Deterministic

citation grounding 0%, nothing is cited

Judge

causal accuracy capped, the direction is inverted

Reviewer

refused sign-off. Contradictions: "Brent fell to $60" (evidence says it rose to $95) and "lower fuel costs boosted margins" (evidence says higher costs pressure them). Fabrications: "$4B buyback" and "Lufthansa alliance", neither of which appears in the evidence.

That is the division of labour. The fabricated answer still reads fluently, and a style-focused grade could be fooled. The reviewer's claim-by-claim fact-check catches the inverted direction and the invented facts and refuses to put a name on it.


Tracing a weakness back to its cause

Grading the finished answer tells you whether Clarity got it right. To make Clarity better, you need to know where it went wrong, because an answer can fall short for three different reasons, each fixed in a different place:

  • Extraction never pulled the fact out of the document, so it was never in memory to begin with.

  • Retrieval had the fact in memory but didn't surface it for this question.

  • Synthesis had the fact in hand but left it out of the written answer.

The evaluation can attribute each gap to the right stage. It re-runs the real extraction on the source documents (instead of using the frozen, pre-extracted version), grades how faithfully extraction reproduced the facts a document contains, then builds memory from that real extraction, runs the agent, and checks, for every concept an answer missed, whether the supporting fact was ever extracted. A missing concept whose fact never made it into memory is an extraction problem. One that did is a retrieval or synthesis problem.

This turns "the answer was thin" into "extraction is dropping the link between these two entities in this kind of document", a precise, actionable target. The fix flows upstream: improve extraction, and every downstream answer that depended on those facts improves with it.

Two levels of relationship grading

A single relationship score hides two very different failures, so the evaluation reports both:

  • Link recall. Did extraction connect the right two entities in the right direction, ignoring the exact verb? This is the durable signal: it asks whether the connection exists at all.

  • Verb-exact recall. Did it also get the predicate right ("becomes CEO of" vs "appointed CEO of")?

The gap between the two is harmless phrasing variance. A missed link is an entity pair the gold connects but the real extraction never joins in the correct direction. It is a genuinely dropped connection, and that is what is worth fixing. (Precision is deliberately not used here: the gold is a curated must-have set rather than an exhaustive one, so good extraction always produces more than the gold and a precision score would punish thoroughness.)

Worked example: extraction was shredding the causal chain

Run across a sample corpus, link recall surfaced a serious, invisible defect. Take an article that says "Siri is an entirely new version powered by Apple Intelligence." The fact that is the story is a causal link:

Real extraction named both entities perfectly, but stored the link backwards and as a plain structural tie:

The same pattern repeated across documents. A regulation that delayed a product was instead filed as "the company announced a delay". A component that powered a platform was filed as "the platform contains the component." Entity recall stayed near 80% the whole time, so an entity-only or verb-only grade saw nothing wrong. Only the orientation-respecting link metric exposed it. Extraction was reliably naming the players while inverting and flattening the causal arrows between them, which is fatal for an engine whose whole job is to explain why.

The fix was upstream, in how extraction is instructed to capture cause and effect. After it, the same corpus moved decisively:

The two previously-missing answer concepts came back at no extra cost, because the causal links that carried them were finally in the graph and pointing the right way. That is the loop working as intended: a weakness the answer-only grade could never localise, traced to a single upstream cause, fixed once, and paid back across every downstream answer.

The remaining link misses after the fix are entity-granularity differences. The gold names a single umbrella concept ("business services") where real extraction split it into the specific services it bundles. That is a definitional difference rather than a dropped connection, which is why link recall and the gap attribution, not a single headline number, are read as the verdict.


What the scores are for

The numbers are a release gate. The point is movement: a retrieval, synthesis, specialization, or consolidation change is run against the evaluation before and after, and it must hold or raise the scores across concept coverage and grounding, the ten judge dimensions, the reviewer's accuracy, reasoning, grounding, and sign-off, and the hardening graders. A change that lifts fluency but drops grounding, or that improves one specialization while regressing another, is caught here rather than in production.

Quality is tracked as a per-surface scorecard across roughly nine investor surfaces — single-stock scenarios, the portfolio card on a demo book, personalized investor-journey personas across their understand, support, and act phases, and thesis proposals — each held or raised independently, so a fix to one surface can't silently regress another. Personalized answers are graded on two extra grounds: their suitability against the investor's profile (goals, risk appetite, horizon, values) and any holdings figures against the live book, while market facts stay grounded only in retrieved evidence.

Because the corpus is fixed and the grading is repeatable, the scores are a durable baseline: the same inputs grade today's agent and tomorrow's, so progress (and regression) are visible run over run.

Last updated