The strength assets of a website and the entity behind it: the indexes, datasets, calculators, downloads, research and reference structures that give a site substance worth citing — plus the external validation that makes engines willing to believe it. The things a brand owns that competitors cannot copy overnight, and that AI engines, journalists and other sites reference as a source. This is deliberately not structured data: schema is the packaging, these are the substance underneath it.
Is there real substance here — data, method, evidence — or a prettier restatement?
Can the engine's crawler actually reach and index the answer at all?
Is the publisher machine-verifiably real — company number, regulator, register links?
Does anyone else say this brand and author are credible? Engines read consensus.
Is it dated, year-tagged and genuinely maintained? Undated loses to dated.
Does the answer exist as a self-contained static passage a machine can lift?
Citation is one payoff of six, not the definition of value. Every asset type now carries a primary job, and the governing rule is: an asset is judged, and cancelled, only against its primary job. An asset that fails at its job is cut; an asset that succeeds at it is never cut for performing poorly at a job it was not built for.
The workhorses. Evergreen, high citation value, and the hardest class for competitors to replicate once established.
Answers: "Who holds the definitive organised view of this niche?"
Numbers you own, or present better than the body that publishes them.
Answers: "Where does the canonical number for this live?"
Tools that answer a personal question. High engagement, strong snippets, natural link targets — and the highest lead capture in the taxonomy.
Answers: "What does this mean for MY numbers?"
Things people take away and keep. Consistently the thinnest class on most sites — which makes it a reliable quick win.
Answers: "What can I walk away with right now?"
Content shaped as an asset rather than an article.
Answers: "Is this a maintained resource or just another blog post?"
Assets whose primary audience is a machine: crawler, engine or integrator. Moves the site from content to infrastructure.
Answers: "Can other systems build on this?"
Strength signals rather than traffic assets. These make every other class more citable — and on YMYL they gate citation eligibility outright.
Answers: "Who is behind this, how are the numbers made, and when were they last checked?"
The class the engines weigh most. Everything in classes 1–7 is self-asserted; engines resolve the brand and author as entities and test whether the wider web corroborates them before citing. Mandatory for every flagship asset and every YMYL site.
Answers: "Does anyone I already trust vouch for this entity?"
AI crawlers execute no JavaScript, click nothing, fill nothing in, and see aggregate assets only in their default state. Four consequences drive most build decisions:
A calculator without a static companion is "a tool exists here" — never the answer. The companion carries the whole citation load.
Tables are cited at row and cell level. Anything behind load-more, filters or pagination needs its own static URL or it cannot be cited.
PDFs are rarely cited, charts are discarded, widgets render as an empty div, gated content is worth zero. The HTML twin does the work.
Heading-question + answer pairs are the cleanest extraction of any format. Timelines win date-shaped queries. Evidence pages prove experience.
The full 18-row reference table — every asset type, what the engine sees, what it cites the page as — is in the Full Framework tab.
Decay is channel-specific and always paired with refresh cost — when an asset rots matters less than what it costs to stop it. A pipeline-fed tracker (fast decay, near-zero cost) is a build; a manually-maintained fast-decay guide is a consolidation candidate.
Never build the full taxonomy on one site. Every candidate asset runs through seven filters, in order — fail one, and it's redesigned, re-scoped or not built.
Can this entity plausibly be cited for this query class? Credentials are half; corroboration is the other half. Self-asserted credibility does not pass.
Does the claimed source or angle genuinely exist for this vertical? "Built on official data" collapses where the dataset doesn't exist.
What does this reveal that current top sources and AI answers do not? If the gain can't be stated in a sentence, it's redundant before it's built.
Evergreen heads → indexes and hubs. Situational long-tail → Q&A and checkers. Annual spikes → refresh assets. Some of the best citation wins have no keyword volume at all.
Document the 5–10 sub-questions engines generate from the head query; confirm each is answered as a self-contained, heading-matched passage. A build requirement.
Incumbents' pattern: thin lists, gated PDFs, undated data. Whole-class white space — timelines, models, case studies, widgets — is the cheapest differentiation there is.
Name the target channel before build: tables → Bing/Copilot, companions → snippets and AI answers, datasets + DOIs → journalists, transcripts → video surfaces, trackers → repeat visits.
The bar is not "new data". The bar is a question with real demand that nobody currently owns the answer to, answered with a method you can defend, and distributed so third parties corroborate it.
Generate data that doesn't exist: operational data, surveys, FOI requests, systematic collection. Highest effort — the only route that makes you the primary source outright.
Cut public data by a dimension nobody has published: region, sector, size, time. The fastest route to ownable research. The cut must reveal, not merely subdivide.
Join datasets never put together: insolvency × payment terms, price × effective dose. The joined view is the original work.
The seven steps: find the unowned question → choose the route → verify nobody owns it → source & document (evidence trail captured during the work) → publish in citation format (capsule → tables → method, CSV + licence, DOI, named author) → distribute (journalist data pack, trade press, dataset deposits — third-party pickup is the point) → build the refresh cycle (annual, year-tagged, versioned; compounding begins at the second release).
Capsule above the tool, tool above the explanation, every sub-question a self-contained passage. Never bury the answer.
Every figure displays its method — and the displayed method is the one actually used. On money and health, anything else is disqualifying.
Source, date and context beside every material claim. Year in the H1 on annual assets. Fake date-refreshing is maintenance theatre and engines detect it.
Every asset carries a named author linked to their hub; YMYL credentials externally verifiable and deep-linked. Voices, not just pages.
Nothing between a tool and its result. The conversion bridge comes after value: post-answer capture, take-away gates, a defined lead path from every high-intent asset.
Owner + refresh trigger + decommissioning rule, or it doesn't ship. An asset without a maintenance plan is a liability with a launch date.
Monthly tracking per flagship across Google AI surfaces, ChatGPT, Perplexity, Copilot and Claude. Listed-among-sources vs cited-for-the-key-claim is the metric that matters.
AI-referred vs organic conversion, assisted conversions via later brand search, lead quality fed back into the roadmap — so citation volume is never mistaken for commercial value.
The framework is a loop, not a production line. Each asset is periodically refreshed, consolidated or retired on the evidence — and the whole thing is applied per brand as a gap map: score the taxonomy against what each site holds, find the white space, filter the candidates, sequence by the ratings layer, attach Class 8 work to every flagship.
Authority Assets are the strength assets of a website and the entity behind it: the indexes, datasets, calculators, downloads, research and reference structures that give a site substance worth citing, plus the external validation that makes engines willing to believe it. They are the things a brand owns that competitors cannot copy overnight and that AI engines, journalists and other sites reference as a source.
This is deliberately not called structured data. In SEO usage, structured data means schema markup (JSON-LD). Schema is the packaging and is covered separately in our technical documentation. Authority Assets are the substance underneath the packaging. Schema makes an asset machine-readable; it cannot make a weak asset strong.
The governing principle of this version: citation is not a property of the asset alone. It is the product of asset quality, retrieval eligibility, entity confidence, external corroboration, freshness and passage extractability. A framework that only manufactures useful pages manufactures content. This framework manufactures a source.
The framework has six parts: the taxonomy of asset classes, the ratings layer, the selection logic, the original research methodology, the build standards, and the measurement and lifecycle layer. It connects directly to the Five Signal Model: Identity Clarity (the entity layer), Subject Authority (assets prove you know the subject), Meaning Architecture (assets structure the subject), Ecosystem Validation (the external layer) and Signal Consistency (dates, sources and named authors that match everywhere the brand appears).
Version 1.0 was reviewed against the current source-selection behaviour of Google organic, AI Overviews and AI Mode, Gemini, ChatGPT, Perplexity, Bing and Copilot, and Claude. Five independent assessments converged on the same finding: the original framework built citable pages but not a believable source. This version adds an eighth asset class (external validation and entity assets), an information-gain filter, query fan-out mapping, static companions for interactive tools, a conversion bridge, corrected schema and llms.txt guidance, anti-scale safeguards, and a measurement lifecycle.
v1.2 rebuilt the ratings layer after a second seven-engine stress test: retrieval and citation are now scored separately per engine, extraction and rendering risk are explicit dimensions, decay is paired with refresh cost per channel, a mechanical reference table records what each engine actually sees per asset type and what it cites the page as, and Classes 7 and 8 moved from enablers to sequencing prerequisites. The ratings are framed openly as working hypotheses: nobody outside the engine teams knows the real weights, so these are best guesses held to account by the Part 6 measurement loop.
v1.3 corrects the citation-lens error the v1.2 restructure introduced: rating everything by citability made "poor citation surface" read as "poor asset", and nearly cancelled legitimate link, ranking and conversion assets. Every asset type now carries a primary job label (RANK, EARN LINKS, RETAIN, CONVERT, GET CITED, SUPPORT TRUST, INFRASTRUCTURE) and is judged only against that job; behind the label sit six payoff columns and four viability modifiers (distribution dependency, uniqueness threshold, maintenance burden, harm exposure). Verdicts corrected on re-audit: spreadsheets and models, embeddable widgets, quizzes, hand-built segment pages, explainers and calendar-driven expert commentary rescued for their non-citation jobs; llms.txt, un-adopted APIs, literal data mirrors, consumer-less feeds and generic curated libraries confirmed weak on every payoff.
Every Authority Asset falls into one of eight classes. A site does not need every type; it needs the right types for its niche, built to standard, plus the eighth class in all cases, because without it the other seven under-perform.
The workhorses. Evergreen, high citation value, and the hardest class for competitors to replicate once established.
Numbers you own, or present better than the body that publishes them.
Tools that answer a personal question. High engagement, strong snippet performance, natural link targets. One rule governs the whole class: the engine cannot operate the tool. No AI crawler executes JavaScript, moves sliders or reads a computed output. The page carrying the tool gets cited as "a tool exists here"; the page that also shows the tool's output in static HTML gets cited as the answer. Every interactive asset therefore ships with a static companion: the formula, worked examples at realistic values, and a results table across common scenarios, in plain server-rendered HTML beneath the tool. This is the machine-readable form of "show the working".
Things people take away and keep. Consistently the thinnest class on most sites, which makes it a reliable quick win.
Content shaped as an asset rather than an article.
Assets whose primary audience is a machine: crawler, engine or integrator.
Strength signals rather than traffic assets. These make every other class more citable.
The class v1.0 lacked, and the one the engines weigh most. Everything in classes 1 to 7 is self-asserted. Engines resolve the brand and author as entities and test whether the wider web corroborates them before citing. This class is mandatory for every flagship asset and every YMYL site.
One structural pattern sits underneath the taxonomy rather than inside it: answer capsules with fan-out coverage. Every asset page opens with a short, self-contained answer block, and the page then covers the natural sub-questions engines generate from the head query, each as its own extractable, heading-matched passage with its supporting evidence close by. Engines assemble answers from multiple sub-query passages, not from one heroic page; passage-level self-containment across the cluster is where citation is won. Capsules are written per page, not templated across pages, because repeated capsule text reads as duplication.
Interactive lead funnels (stepped quote and eligibility wizards) remain commercial assets rather than citation assets, but the wall between the two is a bridge, not a moat: see the conversion bridge standard in Part 5.
Nobody outside the engine teams knows the actual selection weights. Everything in this layer is a working hypothesis, inferred from vendor documentation, observed behaviour and third-party studies, all of which shift without notice. Specific published figures (citation percentages, decay windows, index dependencies) are treated as directional, never as fact. The ratings exist to make our best current guess explicit and testable; the measurement loop in Part 6 is what corrects them. Nothing is discarded permanently on a guess, and nothing is trusted permanently on one either.
v1.2 restructured the ratings around the seven-engine stress test: retrieval and citation are separate events, scored per engine, and the citable surface of every asset is its static HTML text layer. v1.3 corrects the error that restructure created. Rating everything through the citation lens made "poor citation surface" read as "poor asset", and it nearly cancelled legitimate link, ranking and conversion assets, an embeddable widget among them. Citation is one payoff of six, not the definition of value.
Every asset type now carries a primary job, and the governing rule is: an asset is judged, and cancelled, only against its primary job. An asset that fails at its primary job is cut; an asset that succeeds at it is never cut for performing poorly at a job it was not built for. The labels:
Behind the label, each flagship is scored on six payoff columns: organic acquisition, citation acquisition, link acquisition, retention and distribution, lead contribution, and authority support. Four viability modifiers then adjust the verdict: distribution dependency (some assets only work with outreach or a named consumer attached — a widget without a distribution campaign is an expensive calculator, an API without a user is maintenance), uniqueness threshold (how much original substance the type needs before it stops being interchangeable), maintenance burden, and harm exposure (scaled-content, link-scheme, legal and stale-data risk).
AI crawlers execute no JavaScript, click nothing, fill nothing in, and see aggregate assets only in their default state. The consequences per asset type:
| Asset type | What the engine actually sees | What it cites the page as |
|---|---|---|
| Calculator, no companion | Tool shell and surrounding copy; no computed output | "A tool exists here" — never the answer |
| Calculator, with companion | Shell plus formula, worked examples, results table in HTML | The answer for scenario X — citation-ready |
| HTML table (server-rendered) | Rows, cells and headers in the DOM | A comparison or data source — cited at row and cell level, not as a whole table |
| Paginated index | Page 1 plus href links; nothing behind load-more buttons | An incomplete directory, unless a view-all or per-page static URL exists |
| Filterable directory | The default unfiltered state only; filter parameters ignored | A static list; filtered views need their own static URLs (e.g. /directory/london) to exist at all |
| Linear text, layout often jumbled, images and charts lost | Rarely cited; the HTML twin is the citation surface | |
| Spreadsheet / model | Download link and description; cell contents unreachable | "A model exists"; the methodology page earns the citation |
| Dataset (CSV/JSON) | The file plus its explanatory landing page | The landing page in search contexts; research modes read the file itself and cite heavily |
| Video / audio | Player, metadata, and the transcript if present in HTML | The transcript; without one, only "a video exists" |
| Chart / infographic | An img tag, alt text and adjacent caption; the visual data is discarded | The surrounding prose and data table; the visual itself is invisible |
| Embeddable widget | An empty div or script call; nothing renders | Nothing — backlink value only, and that is the point of building one |
| API | The documentation page; endpoints are never executed | The documentation, as infrastructure proof |
| Structured feed | Entries and timestamps | Not cited; discovery and freshness plumbing (strongest in the Bing/Copilot pipeline) |
| Gated content | The gate and the teaser copy | Nothing; zero citation value behind any gate — capture is its job, not visibility |
| Q&A page | Heading-question and answer pairs; the cleanest extraction of any format | The answer to the specific question |
| Glossary entry | Term-definition pairs | A definition source for niche terms only; institutions own generic definitions |
| Timeline | Date-event pairs | A historical fact source; clean extraction, strong on date-shaped queries |
| Evidence page | First-hand text, logs and described media | Proof of experience — increasingly a primary selection signal, not just trust decoration |
Decay is channel-specific, and the layer now pairs it with refresh cost, because when an asset rots matters less than what it costs to stop it. A pipeline-fed tracker (fast decay, near-zero refresh cost) is a build; a manually-maintained regulatory guide (fast decay, high refresh cost) is a consolidation candidate before it is a build.
| Channel | Freshness sensitivity | Working notes |
|---|---|---|
| Perplexity | Highest | Real-time retrieval; reported to punish data pages with no substantive modification hardest. Trackers and stat pages need the most aggressive refresh here. |
| Google AI Overviews / AI Mode | High | Dated, year-tagged sources favoured; undated data assets are treated as stale on YMYL queries. |
| Bing / Copilot | Medium-high | IndexNow-driven; year-tagged content and freshness signalling weigh heavily, and updates propagate fastest of any pipeline. |
| Gemini | Medium-high | Temporal signals weighted, but entity confidence appears to dominate: corroborated identity beats recency. |
| Google organic | Medium | Buffered by links and history; evergreen content survives here longer than anywhere else. |
| ChatGPT | Variable | Live search behaves like search; cached and trained retrieval barely weighs freshness, where pre-cutoff entrenchment matters more. Treat as two channels. |
Each type carries its primary job, effort, an honest payoff description (citation and non-citation), and decay paired with refresh cost. Verdict corrections from the non-citation re-audit are reflected throughout.
| Asset type | Job | Effort | Payoff reality | Decay and refresh cost |
|---|---|---|---|---|
| Indexes | GET CITED | High | Retrieved widely; cited at row level only, and only rows the crawler can reach — needs a view-all or static per-segment layer. Also ranks and earns links | Compounds; scheduled refresh, moderate cost |
| Comparison tables | RANK | Medium | Cited cell- and row-level with clear headers, dating and disclosure; undisclosed pay-to-play kills selection | Rots without versioning; low-moderate cost |
| Glossaries | RANK | Low | Citation niche-terms only, but built for long-tail rankings and internal-link architecture at near-zero cost; term pages need genuine search demand | Near-zero decay |
| Directories | RANK | Medium | Strongest in Bing/Copilot; YMYL needs verifiable inclusion rules; filter states invisible without static URLs; leads follow position | Rots as providers change; moderate cost |
| Timelines / chronologies | GET CITED | Low-med | Cleanly extracted; strong on date- and regulation-shaped queries | Append-only; low cost |
| Curated libraries | RETAIN | Low | Rarely cited; worth building only where governed curation solves a real research task (maintained, dated, inclusion rules) or the "best books on X" query genuinely exists. Otherwise skip | Maintenance is the whole asset |
| Asset type | Job | Effort | Payoff reality | Decay and refresh cost |
|---|---|---|---|---|
| Stat pages / roundups | GET CITED | Low-med | Strong: passage-dense, source-adjacent claims are what answer engines lift | Source-release cycle; low-moderate cost |
| Original research | GET CITED | High | Highest citation value on every channel; converts only with Class 8 distribution attached; also the fleet's best link earner | Compounds if versioned; annual cost |
| Trackers | GET CITED | Medium | Strong where genuinely fresh; a stale tracker is worse than none in AI channels; also earns repeat visits and feeds widgets | Fastest decay of any type; near-zero cost pipeline-fed, high cost manual |
| Benchmarks | GET CITED | Medium | Cited as the measuring stick when the method is visible; strong conversion adjacency | Annual refresh; moderate cost |
| Public data mirrors | RANK | Low-med | Literal mirrors stay skip. Transformed mirrors — a usable, filterable, historical layer over a hostile official source — rank, win the click and earn journalist links even where the AI answer credits the official body | Source cycle; pipeline cost is real |
| Annual-refresh assets | GET CITED | Low-med | Year-tagged queries favour them across AI channels; versioned editions compound | Value is the discipline; low cost |
| Evidence assets | SUPPORT TRUST | Varies | A primary selection signal for experience-weighted queries, not trust decoration | Permanent record; minimal cost |
| Asset type | Job | Effort | Payoff reality | Decay and refresh cost |
|---|---|---|---|---|
| Calculators | CONVERT | Medium | An organic destination and lead engine; every AI channel cites the static companion only | Rots on rate change; cost tracks rule changes |
| Checkers / eligibility | CONVERT | Medium | The stated rules get cited; the companion carries the citation load; the tool qualifies the lead | Rots with criteria changes; moderate cost |
| Quizzes / diagnostics | CONVERT | Low-med | Rarely cited and that is fine: built as triage and qualification — route the user by situation, capture after value, pass context into the conversion journey. Result pages shareable | Slow decay; compliance review cost in YMYL |
| Configurators / estimators | CONVERT | High | Scenario and companion pages cited; the tool itself is invisible to every engine | Holds; high build cost |
| Asset type | Job | Effort | Payoff reality | Decay and refresh cost |
|---|---|---|---|---|
| Templates | CONVERT | Low | Page cited for existence; content cited only where mirrored in HTML; strong capture | Near-zero decay |
| Checklists | RANK | Low | Structured-list extraction is snippet-shaped; cheap long-tail wins | Slow decay |
| Guides (HTML primary, PDF optional) | CONVERT | Medium | The HTML is the entire citation and ranking surface; the gated PDF is the capture trade and adds nothing to visibility — which is fine, capture is its job | Regulation cycle; generate PDF from the HTML source |
| Spreadsheets / models | CONVERT | Medium | Raised: a working file is a work product, not content — practitioner links, branded circulation, and the most legitimate capture gate in the taxonomy (users enter real numbers). Methodology page carries the citation. Formula errors are reputationally expensive | Slow decay; version every release |
| Sample documents | RANK | Low | Cited when the scenario matches the query | Slow decay |
| Datasets (CSV/JSON) | GET CITED | Low | Heavily cited in research modes (Perplexity, deep research); the landing page carries the search citation | Tied to parent asset; low cost |
| Asset type | Job | Effort | Payoff reality | Decay and refresh cost |
|---|---|---|---|---|
| Pillar hubs + spokes | RANK | High | Organic architecture; answer engines cite spokes and passages, never the hub itself | Compounds; moderate cost |
| Q&A libraries | GET CITED | Medium | Raised for snippets and AI answers when built from real demand; templated versions read as scaled content and are suppressed | Slow decay; add on demand |
| How-to / process guides | RANK | Medium | Cited at individual step level | Rots with process change; moderate cost |
| Explainers | RANK | Low | Re-rated: citation is niche-only but that was never the job — explainers are the long-tail ranking spokes and internal-link targets that feed every hub, and they assist conversion for unfamiliar users | Slow decay |
| Case studies | CONVERT | Medium | Cited when the scenario matches; doubles as experience proof and conversion confidence | Near-zero decay |
| Expert commentary streams | EARN LINKS | Ongoing | Rarely cited as fact and not built to be: built for reactive journalist quotes on the official release calendar (monthly stats days, rate decisions), compounding the author entity. Requires a named expert and a distribution process, or skip — a stream that stops visibly signals abandonment | Compounds author; ongoing cadence cost |
| Troubleshooting clusters | RANK | Medium | Distress queries are snippet-shaped; high-anxiety extraction wins, and conversion follows | Slow decay |
| Segment / matrix pages | RANK | Medium | Re-rated: hand-researched pages for real, commercially meaningful segments (sector, borrower type, situation) rank and convert strongly — the risk was never segment pages, it is the templated matrix. Hard scale cap; each page passes the uniqueness threshold or is not built; enforcement risk stays priced in | Rots with criteria; high maintenance cost |
| Multimedia (with transcripts) | RETAIN | Med-high | The transcript is the citable layer and reaches video-weighted surfaces text never touches; the video itself builds brand and repeat use | Slow decay with transcripts |
| Asset type | Job | Effort | Payoff reality | Decay and refresh cost |
|---|---|---|---|---|
| APIs | EARN LINKS | High | The documentation is cited, never the endpoint. Payoff arrives only with adoption, so the rule is promote before build: push existing APIs at journalists and tool-builders; commission nothing new until one shows uptake. An unused API is maintenance | Compounds only with adopters |
| Embeddable widgets | EARN LINKS | Medium | Rescued: never a citation surface, and that was never the job. A pipeline-fed data widget that saves another publisher work earns in-content topical links. Branded or nofollow attribution only (keyword-anchor widget links are link spam), and a widget without a distribution campaign is an expensive calculator. Every embed is a forever-maintenance promise | Holds only if maintained; high risk |
| Structured feeds | INFRASTRUCTURE | Low | Retrieval and freshness plumbing (strongest in the Bing/Copilot pipeline); add to pipeline-fed data pages because it costs nothing; build for no other reason unless a named consumer exists | Holds; near-zero cost |
| Code repositories | SUPPORT TRUST | Low-med | Reproducibility proof beside published research; technical citation elsewhere. A neglected repo is negative evidence | Slow decay |
| llms.txt / AI instructions | None (hygiene) | Low | Fails all payoffs; kept as cheap hygiene under strict standards (factual only, no instructions, verified), never in competition with real assets | Verify on schedule |
| DOI publications | SUPPORT TRUST | Medium | Formally citable; earns academic-domain links and entity legitimacy, only with Class 8 distribution behind it. Never create research merely to obtain a DOI | Permanent by design |
| Asset type | Job | Effort | Payoff reality | Decay and refresh cost |
|---|---|---|---|---|
| Methodology / policies / publisher identity | SUPPORT TRUST | Low | Prerequisite, not enabler: gates YMYL citation eligibility outright | Holds with review dates |
| Data Sources & Updates page | SUPPORT TRUST | Low | Prerequisite; the freshness proof engines and users both check | Value is the update discipline |
| Author hubs + entity package | SUPPORT TRUST | Low-med | Prerequisite; entity resolution weighs heavily everywhere and appears to dominate in Gemini | Compounds per publication |
| Versioned knowledge products | SUPPORT TRUST | Low | Freshness and correction proof; research modes cite change logs directly | Append-only |
| Third-party corroboration / PR | EARN LINKS | Ongoing | The selection layer itself; converts the citation rate of everything above | Compounds |
| Distribution assets | EARN LINKS | Low-med | Multiplier on research and data assets | Holds |
Before a flagship asset is commissioned: first, name its primary job and the success metric that job implies — the asset will be judged against that and nothing else. Then apply two hard gates. Gate one, rendering: if the asset's answer only exists behind interaction, JavaScript, a gate or a non-HTML format, it fails for every AI channel until the static layer is designed — unless GET CITED is not its job. Gate two, authority anchor: on YMYL query classes, if the publisher entity is not machine-verifiably linked to its anchors (Companies House, regulator register, ICO), citation likelihood is capped regardless of content quality. Then score the six payoff columns (organic, citation, links, retention, leads, authority support), scoring retrieval and citation separately per engine (Google organic, AI Overviews and AI Mode, Gemini, ChatGPT live and cached, Perplexity, Bing/Copilot, Claude); then apply the four viability modifiers: distribution dependency, uniqueness threshold, maintenance burden, harm exposure. An asset scoring low on its primary job is redesigned or dropped before build; an asset scoring low elsewhere is not.
Reading the ratings: indexes, trackers, calculators-with-companions and original research remain the four disproportionate-strength types for citation, each earning it only through its static, extractable layer. The v1.3 re-audit rescued a second tier whose jobs are not citation: spreadsheets and models (CONVERT), pipeline-fed embeddable widgets (EARN LINKS), quizzes and triage diagnostics (CONVERT), hand-built segment pages (RANK) and calendar-driven expert commentary (EARN LINKS). Classes 7 and 8 remain prerequisites that gate citation eligibility itself, and they sequence first.
Never build the full taxonomy on one site. Run every candidate asset through seven filters, in order.
Can this entity plausibly be cited for this query class? For YMYL queries, engines select government sources, regulators, institutions and externally corroborated experts, and deeper reasoning modes weight institutional sources more heavily still. Credentials answer half the question; the other half is corroboration: who else says this brand and author are credible, and does the wider web reflect it? If the ceiling test fails, change the query class, close the corroboration gap through Class 8 work, or do not build the asset. Self-asserted credibility does not pass this filter.
Does the claimed source or angle genuinely exist for this vertical? "Built on official data" works where an official dataset exists and collapses where it does not. A positioning hook that only fits part of a portfolio must not be stretched across all of it.
What does this asset reveal that the current top sources and AI answers do not? A new observation, cut, series, calculation, comparison dimension or decision rule counts; a prettier restatement does not. The unique contribution must be identifiable in one or two extractable passages, and provably derived rather than rewritten. If the gain cannot be stated in a sentence, the asset is redundant before it is built.
Is there real demand, and what shape is it? Evergreen head terms suit indexes and hubs; situational long-tail suits Q&A and checkers; annual spikes suit refresh assets; seasonal and visual demand adds Discover and multimedia angles; news events suit commentary streams. Measured keyword volume is not the whole picture: some of the best citation opportunities recur across thousands of varied conversational prompts with no conventional volume, and mid-funnel comparison intent often converts best of all. Map the query journey (discovery, comparison, validation, action), not just the keyword.
Engines split the user's question into sub-queries and cite pages that surface consistently across the set; ranking first for the headline term no longer guarantees selection. For every asset, document the natural sub-questions (typically five to ten) that fan out from the target query, and confirm the asset or its cluster answers each one as a self-contained, heading-matched passage. Fan-out coverage is a build requirement, not an editorial nicety.
Where are incumbents thin? The recurring pattern: thin provider lists, gated PDFs, undated data. The counter: expert-updated depth, un-gated tools, shown working, dated sources. Whole-class white space (timelines, models, case studies, widgets are commonly absent everywhere) is the cheapest differentiation available.
Name the target channel before build. Utility tables surface in Bing, which feeds Copilot; capsules and companions win snippets and AI answers; datasets and DOIs earn journalist and academic citation; video and transcripts reach YouTube-weighted surfaces; trackers earn repeat direct visits; community presence earns experience-query citations. Register the site on the channels the assets need: Search Console, Bing Webmaster with IndexNow, and explicit crawler access per engine (OAI-SearchBot, GPTBot, PerplexityBot, Claude, Google-Extended) as deliberate policy decisions, not defaults.
For a new or thin site: trust template, publisher identity and Data Sources & Updates page first, with the entity package (Class 8 schema, register links and verifiable publisher identity) built first, because Classes 7 and 8 are prerequisites rather than enablers: on the current evidence, identity and external validation gate citation eligibility itself, especially on YMYL queries. Then one flagship index or tracker for the niche. Then the cheap wins: glossary, templates, CSV downloads of data already held. Then the Q&A library built from real demand. Original research ships last, and never with only self-published credibility: every flagship asset launches with a distribution plan and a named list of who will corroborate, use or cite it. Research published before the trust and entity layers exist is wasted; nobody can verify who stands behind it.
Original research is the highest-value asset class and the least understood. The bar is not "new data". The bar is a question with real demand that nobody currently owns the answer to, answered with a method you can defend, and distributed so that third parties corroborate it. There are three routes.
Generate data that does not exist: your own operational data, a commissioned or self-run survey (with stated sampling and consent standards), freedom of information requests, or systematic collection from public records. Highest effort, highest defensibility, and the only route that makes you the primary source outright. The question comes first, the collection method second; never collect data and then hunt for a question.
Take existing public data and cut it by a dimension nobody else has published: region, sector, company size, time of year. A national statistic becomes a regional index; a UK-wide rate becomes a sector league table. The insight is new resolution, and the specific cut exists only on your page. The fastest route to ownable research and the default for most sites. The information-gain test applies in full: the cut must reveal something, not merely subdivide.
Join two or more public datasets that have never been put together: insolvency rates against payment terms by sector, adviser density against regional demand, price against effective dose. The joined view is the original work. Recombination also powers combination tools, where the user supplies one dataset themselves.
Topic given: invoice finance. The commodity content is saturated and earns nothing. The research process:
Every Authority Asset, in every class, meets these standards before it ships. Most were learned from observed failures; each names the failure it prevents.
v1.0 treated visibility as a one-way output of asset quality. It is not: retrieval pipelines differ per engine and change without notice, and no single metric exists. Every flagship asset therefore carries a measurement loop.
The framework is applied per brand as a gap map: the taxonomy is scored against what each site already holds, white space identified (fleet-wide white space first, since it differentiates every site at once), candidates run through the seven filters, flagship assets scored on the channel scorecard, and the build list sequenced by the ratings layer with Class 8 work attached to every flagship. The applied view for each brand is maintained as a live companion to this document.