We scored 5,892 articles. 60.6% fell below the bar — and 93.8% of those were hitting a ceiling, not the floor.
A meta-analysis of the article-quality distribution behind ANIA. The uncomfortable finding is not how much of our corpus scored below the activation threshold. It is how little of that was actually bad.
The first number a content-quality report hands you is the scary one: the share of your corpus that scored below the bar. Ours was 60.6%. Read naively, that says more than half the articles ANIA has ever analysed are not good enough to surface. That is the number that makes you want to tear the pipeline down and start again. It is also, almost entirely, the wrong number.
This post is the same kind of read as our field notes on failures disguised as success and the escalation boundary: aggregate the data, then ask which single number is hiding the real shape. Every figure below is a live row from the same content-quality deepdive that scores our corpus, frozen to a single cut so it reproduces.
What we aggregated
ANIA scores every ingested article on a 0–100 relevance scale; anything at or above 75 becomes “active” and is eligible for newsletters and topic pages. We took the all-time scored corpus — 5,892 articles — and binned every score into the histogram the scoring pipeline already uses.
The shape of that histogram is the first clue something is off. A healthy scoring system produces a spread — a tail of weak material, a body of middling scores, a cluster of strong material near the top. Ours does not. It has a wall at 75: an enormous pile of articles wedged into the five-point band immediately below the threshold, and comparatively little anywhere else below it.
The reframe: sub-threshold is two populations, not one
“How many articles are below the bar” treats everything under 75 as one undifferentiated mass. It is not. There are articles that scored badly because they are bad, and articles that scored just-below the bar because the scoring rubric structurally caps them there. Mixing the two is the defect: the second group inflates the first and makes a calibration problem look like a content crisis.
That table is the whole post. Of the 3,573 articles below the activation threshold, 93.8% — 3,350 of them — are near-misses parked in the band a hairsbreadth under the bar. Only 223 articles (6.2% of the below-bar population) are genuinely low quality. The scary 60.6% figure is, almost in its entirety, a measurement ceiling.
Why a ceiling, not a floor
The analyst rubric is deliberately conservative about non-regulatory, quantitative business news: unless an article carries a primary-source catalyst we can attribute, it is capped below the activation threshold regardless of how well-written it is. That cap is a feature — it keeps thinly-sourced market chatter out of the signal. But it has a side effect: a large volume of competent, on-topic articles score exactly in the 70–75 band, because that is where the cap puts them. They are not failing the quality bar. They are failing the attribution bar, which is a different bar with a different fix.
This is why the raw “sub-threshold count” is an actionable signal only after you split it. A near-miss-heavy distribution tells you your scoring is calibration-bound — the work is either widening the attribution path or accepting that the ceiling is intentional. A genuinely-low-heavy distribution tells you something is wrong with your sources or your ingestion. Those are opposite diagnoses, and the aggregate number hands you neither.
The lesson that generalises
We keep arriving at the same shape of finding across very different parts of the system. A run-level “success” that did nothing; an escalation that never reached a human; a below-bar count that is mostly a ceiling. The shared property: the headline aggregate is true and misleading at the same time, and only the disaggregated view tells you whether you have a quality problem or a measurement problem. The disaggregation is the whole job.
The measure of a trustworthy quality report is not how low the low score is. It is whether the number you are reacting to describes the thing you think it does — and “how many articles are bad” is a question you cannot answer until you have separated the ceiling from the floor.
The ANIA promise
Quality is only real where the number is honest.
We build ANIA in the open — including the meta-analyses of where our own scoring quietly misleads us. If field notes on the honest edges of an intelligence product are useful to you, follow along.
Try ANIA free