Grammar Templates

Answers that matchthe standards ofresearch.

Institutional finance expects a specific register: conclusion-first, figures woven into prose, calibrated hedging. We distill that register automatically from ~12,800 research documents into a single system-prompt style template, then measure it in blinded, open-book A/B tests. It wins decisively; and the gains are substantive, not cosmetic.

90.5%
Preferred over a no-template baseline in the best model, blinded pairwise, held-out judge
+3.9 pts
Key-fact recall rises under the template; the gains are substantive, not just stylistic
0/30
Topics where a specialized per-topic template beat the single global template

Two arms, one variable

Both arms answer the same question from the same source excerpts, open-book, so facts are held constant. The only thing that changes is whether the corpus-derived style template is in the system prompt.

Baseline

No template

A neutral but prose-capable system prompt, explicitly told to write well-structured prose, not bullet lists. So we measure answer quality, not a prose-vs-bullets formatting artifact.

Templated

+ global style template

The same prompt plus a single corpus-derived template: 8 style sections (sentence construction, data integration, hedging, framing…) and a domain layer of key terms, formulas, and relationships.

Does the template help?

Blinded pairwise preference with randomized answer order, judged by a held-out model over decisive (non-tie) pairs. The template wins decisively.

A second Claude judge (Sonnet 4.6) agrees — 89.5% (Opus) and 82.9% (Haiku). And the template doesn't trade accuracy for polish: independently-scored key-fact recall rises under it.
92.2%94.7%+2.6
Key-fact recall — strong model (Opus 4.8)
85.0%89.0%+3.9
Key-fact recall — small model (Haiku 4.5)

Robust across judge vendors

Because the answers and the primary judge are all Claude, self-preference is a real risk. So the same answers were re-scored by three independent non-Anthropic judges. Every judge favors the template.

No consistent same-vendor inflation. On the strong model's answers the Claude judge (90.2%) is mid-pack among the neutral judges (78.9–92.5%); on the weak model's answers it is the most conservative (61.0% vs 80.8–87.9%).

One global template beats specialization

Should each topic get its own template? We built 89 per-topic templates and tested them with oracle routing (the item's true topic, an upper bound that removes routing error). Even so, specialization does not beat a single broad template.

0 of 30 per-topic comparisons were significant after Benjamini–Hochberg correction, including the newly added non-equity topics (credit, commodities, macro/FX). Meanwhile the global template's win over baseline replicates on this independent set (85% / 71%). One broad, domain-enriched template generalizes best.

How it's built

A three-stage pipeline, entirely corpus-derived, no hand-written style rules.

Corpus

~12,800 research documents from all providers, 2024-onward, ~12 regions. Provider-balanced and asset-class-diverse: roughly two-thirds equity, one-third non-equity (credit, commodities, macro/FX).

Per-document analysis

A model reads each document and extracts a topic, region, and a style template (sentence construction, data integration, tone/hedging, etc) plus a layer of key terms, formulas, and relationships.

Synthesis

Per-document templates are merged hierarchically into one global template: 8 style sections plus 3 consolidated domain sections spanning equity, credit, rates/FX, and commodity conventions.

Evaluation

Blinded pairwise preference with randomized order, plus a separate key-fact recall metric. Open-book throughout, so facts are held constant and only the style guidance varies.