We ingested 5,892 articles. Three sources were 90.1% of all of them.
A meta-analysis of source concentration behind ANIA. The uncomfortable finding is not how much of our corpus scored below the bar. It is how little of it came from anywhere other than three RSS feeds.
The first number an ingestion report hands you is a reassuring one: how many sources you pull from. Ours was 13. Read naively, that says ANIA reads across a diversified field of publications. It is also, almost entirely, the wrong number. The question that matters is not how many feeds are wired up. It is how much of what you actually read comes from any one of them.
This post is the same kind of read as our field note on what the quality distribution was hiding and the pipeline debt that quietly destroys trust in an intelligence product: aggregate the data, then ask which single number is hiding the real shape. Every figure below is a live row from the same content-quality deepdive, frozen to a single cut so it reproduces.
What we aggregated
ANIA ingests articles from a configured set of feeds and scores each one on a 0–100 relevance scale. We took the all-time scored corpus — 5,892 articles across 13 distinct sources — and asked a single question economists have asked about markets for decades: how concentrated is this?
The shape of that table is the whole finding. Hacker News alone is 31.3% of every article ANIA has ever analysed. Just two of them — Hacker News and TechCrunch — are 61%. The three largest sources — Hacker News, TechCrunch, and The Verge — together account for 90.1%. The remaining 10 sources split the last 9.9% between them. This is not a corpus. It is three feeds with a tail.
The number that grades it: HHI 2736
“Three sources are most of it” is an intuition. There is a formal version of that intuition, and regulators use it to decide whether a market is a monopoly. It is the Herfindahl-Hirschman Index: square each source’s percentage share, sum them, and you get a single number on a 0–10,000 scale. The higher it is, the more concentrated the market.
The HHI for our corpus is 2736. On the ladder regulators use, anything above 2,500 is graded highly concentrated — the band reserved for markets where a small number of players dominate. By the exact same standard, our “diversified” ingest of 13 sources is, mathematically, a oligopoly of three.
Why concentration is a reliability defect, not a vanity metric
A high HHI is not interesting because it is embarrassing. It is interesting because concentration is a single point of failure wearing a corpus’s clothes. If Hacker News changes its ranking algorithm, tightens its rate limit, or just has a slow week,31.3% of our pipeline moves with it. If two of the top three do, we lose most of what we read. An intelligence product whose coverage swings with the uptime of three websites does not have a sourcing strategy. It has three dependencies it calls a strategy.
There is a second, quieter cost. Concentration is a coverage-bias amplifier. The three feeds at the top of our table are excellent at what they cover — but what they cover is the same waterfront of consumer-tech and venture news. The specialist sources further down the table (the ones at 0.1–0.3% of the corpus) are where the on-topic signal for a niche intelligence product actually lives. A corpus that is 90.1% top-three feeds by volume will, by construction, surface top-three-feed perspectives. The long tail is where relevance is; the head is where the volume is; and a naive pipeline conflates the two.
The lesson that generalises
We keep arriving at the same shape of finding across very different parts of the system. A run-level “success” that did nothing; a below-bar count that was mostly a scoring ceiling; an escalation that never reached a human. The shared property: the headline aggregate is true and misleading at the same time, and only the disaggregated view tells you whether you have a real problem or a measurement problem. “We pull from 13 sources” is the aggregate. The HHI is the disaggregation.
The measure of a trustworthy ingest is not how many feeds are wired up. It is whether the volume behind those feeds is honest about how much of your coverage depends on a handful of them — and “how diversified is our sourcing” is a question you cannot answer until you have squared the shares and summed them.
The ANIA promise
Coverage is only real where the sourcing is honest.
We build ANIA in the open — including the meta-analyses of where our own ingest quietly over-concentrates. If field notes on the honest edges of an intelligence product are useful to you, follow along.
Try ANIA free