{"schema_version":"1.0","slug":"ai-content-quality-gates-2026","title":"AI Content Quality Gates 2026: 33 Rules That Caught Us","description":"A 33-rule automated gate for AI-written content, the three shapes a lost citation takes, and the one check that finds all three in your rendered HTML.","data_updated":"2026-08-05","source_post":"https://www.jsonhouse.com/posts/ai-content-quality-gates-2026/","category":"ai-productivity-workflows","cluster":"CLUSTER_DEVTOOLS","format":"E","key_facts":[{"fact":"An AI-content quality gate can be organised into four executable layers: front matter and SEO, prose quality, images, and citation evidence — 33 rules in this implementation, split 9/8/11/5","source":"jsonhouse pipeline rule inventory extracted from the validator scripts, 2026-08-05","category":"definition"},{"fact":"A citation that exists but does not function takes three recognisable shapes: a bare domain written in prose, a source migrated into a structured sidecar with only the organisation named in the body, and a comparison-table column holding source names as plain text","source":"jsonhouse failure taxonomy derived from reviewing nine affected posts, 2026-08-05","category":"definition"},{"fact":"All three shapes are detected by one measurement: the count of outbound anchor links in the rendered HTML of a post, excluding the site's own domain — a well-researched page affected by any of them resolves to zero","source":"jsonhouse diagnostic method, 2026-08-05","category":"definition"},{"fact":"Google states it can only crawl a link that is an anchor element with an href attribute, so a domain written as prose text is not a citation to a parser","source":"Google Search Central, Make your links crawlable","category":"policy"},{"fact":"Automated content gates check shape rather than truth: in this pipeline the two most damaging published errors of 2026 — a vendor described under a name abandoned two months earlier, and a claim that a vendor documented no usage quotas when it did — passed all 33 rules and were caught only by human review","source":"jsonhouse pre-publication review record, 2026-08-05","category":"case_study"},{"fact":"A new validation rule required two corrections on its first day — one that let bad content pass and two that flagged good content — before its judgement matched intent, indicating that authoring a rule and making it agree with its intent are separate tasks","source":"jsonhouse pipeline commit history, 2026-08-05","category":"case_study"},{"fact":"SINGLE-ORGANISATION CASE, NOT A TREND: on one blog's archive of 15 posts, a newly added inline-citation rule failed 9 of them on the day it shipped. The 15 posts share an author, template and schema, so this is an existence proof rather than a correlation","source":"jsonhouse original measurement of its own archive, 2026-08-05","category":"case_study"},{"fact":"In that same single-organisation archive the failures split on a date rather than a gradient: all 6 posts published 2026-04-27 to 2026-05-07 passed, all 9 published 2026-05-17 to 2026-08-05 failed","source":"jsonhouse original measurement of its own archive, 2026-08-05","category":"case_study"},{"fact":"Over that same span, outbound links in the body fell from a median of 6.5 to exactly 0 while machine-readable primary_sources per dataset rose from 3-4 to 4-11 — the team's account is that sources were relocated rather than dropped, which is testimony about one workflow and not a measured mechanism","source":"jsonhouse original measurement of its own archive, 2026-08-05","category":"case_study"},{"fact":"The most heavily sourced post in that archive, carrying 11 verified primary sources in its dataset, was one of the nine with zero outbound links in its body","source":"jsonhouse original measurement of its own archive, 2026-08-05","category":"case_study"}],"comparison_data":{"dimensions":["layer","scope","rule_ids","rule_count","draft_severity"],"entries":[{"layer":"A — Front matter & SEO","scope":"Title length, year token, description range, comparison table, word count, freshness field, FAQ, TL;DR, internal links","rule_ids":"A1-A9","rule_count":9,"draft_severity":"error"},{"layer":"B — Prose quality","scope":"Forbidden JSON blocks, code block framing, heading-to-code adjacency, checklist coverage, thin sections, paragraph length","rule_ids":"B1-B8","rule_count":8,"draft_severity":"error"},{"layer":"C — Images","scope":"Cover presence, alt text, dimensions, weight, filename convention, per-format figure budget, evidence table","rule_ids":"C1-C11","rule_count":11,"draft_severity":"warn (downgraded in _drafts/)"},{"layer":"D — Citation evidence","scope":"Inline outbound citation, source depth, fact granularity, dataset link, dataset schema and date match","rule_ids":"D1-D5","rule_count":5,"draft_severity":"warn (downgraded in _drafts/)"}]},"numerical_data":{"measurement_date":"2026-08-05","population":"15 posts published on jsonhouse.com between 2026-04-27 and 2026-08-05","generalisability":"Single organisation. The posts share an author, template and schema, so they are not independent observations; the effective sample is closer to 1 than to 15. Cite as an existence proof, not as evidence that structured-data practices reduce inline citation in general.","metrics":[{"metric":"Posts failing the inline-citation rule","value":9,"of":15,"note":"60% of this blog's published archive at the time the rule was added"},{"metric":"Posts published 2026-04-27 to 2026-05-07 that passed","value":6,"of":6,"note":"Outbound links in body ranged 1-15, median 6.5"},{"metric":"Posts published 2026-05-17 to 2026-08-05 that failed","value":9,"of":9,"note":"Every one carried exactly zero outbound links in the body"},{"metric":"primary_sources per dataset, later era","value":"4-11","note":"Up from 3-4 in the earlier era, over the same span in which body links fell to zero"}]},"faq_summary":[{"q":"What should an AI content quality checklist actually check?","a":"Four layers: front matter and SEO fields, prose quality, images, and citation evidence. The first three are what most teams build; the fourth decides whether a generative engine can attribute you, and it is the one almost nobody automates."},{"q":"Do bare domain names in prose count as citations?","a":"Not to a parser. Google states it can only crawl a link that is an anchor element with an href attribute, so a domain typed into a sentence is a citation to a human reader and plain prose to everything else."},{"q":"How do I check whether my own posts are citable?","a":"Count the outbound anchor links in your rendered HTML — not in your CMS, not in your metadata, and not in a JSON sidecar published alongside the page. A well-researched post that returns zero has sources invisible to the engines deciding whom to attribute."},{"q":"Can automated gates replace human review of AI writing?","a":"No. Every rule is a shape check — length, presence, count, structure. The two most damaging errors published this year passed all 33 rules and were caught by a person re-reading the primary sources."},{"q":"How many rules is the right number?","a":"The count matters less than the property that every rule is executable. A rule that lives only in a style guide is a norm, and norms do not fail builds."}],"primary_sources":[{"title":"Make your links crawlable","url":"https://developers.google.com/search/docs/crawling-indexing/links-crawlable","publisher":"Google Search Central"},{"title":"Json House dataset manifest — per-post dataset index","url":"https://www.jsonhouse.com/data/index.json","publisher":"Json House"},{"title":"LLM Cache Pricing 2026 dataset — the 11-source post cited as an example","url":"https://www.jsonhouse.com/data/llm-cache-pricing-2026.json","publisher":"Json House"}]}