Skip to content
HEXART report · 29 August 2026

224 AI-written articles: our own site report

For five months, we published posts written by a language model on this website. 224 were produced. Exactly 18 percent answer a question that someone actually asks. Below are the numbers, our methodology, and what changed after a single pipeline fix.

We are publishing this because such data does not exist. You can find guides on generating AI content, and warnings about search engine penalties. No company has shown the full record of their own archive, including the parts that failed. This is that case study.

What we measured

The sample consists of 224 materiały publiczne from the hexart.pl news section, published between 3 April and 27 August 2026. The same pipeline generated every post: the model picked the topic, wrote the copy, selected images, and recorded the voiceover. A human approved publication, but wrote no copy.

We split the data into two periods: on 19 August 2026, we updated the pipeline. Before writing, it now searches for questions lacking good answers, and it halts publication if a source list is missing. This is the only point in the data showing the direct impact of a specific fix rather than an overall trend.

Numbers

7
MetricThrough 18.08.2026From 19.08.2026Why it matters
Published posts19311We slowed production deliberately: from 82 posts in April down to one per day. Slowing down alone did not improve quality, as the next rows show.
Title matches an actual search query34 (18%)11 (100%)The most important number here. Four fifths of the output from the first four months answers no search query: asked for "industry news", the model invents a label for a concept and writes around it.
Post includes a source list20 (9%)11 (100%)The research step already existed, but its output was lost before saving. After the fix, a post without sources never goes live. It is not published with lower quality; it is not published at all.
Made-up acronym in parentheses in the title20An extreme case of the same mechanism: the model coins a name for a phenomenon, then shortens it to a three-letter acronym to sound like a technical term.
The article cites an amount, a deadline, or a number0 (0%)10 (91%)Zero is not a metaphor. Across 193 articles from the first period, not a single amount in PLN appears: an article about a non-existent phenomenon has nothing to price.
In-paragraph citation to a source0 (0%)11 (100%)A source list at the bottom is one thing; an inline link next to a claim is another. Without it, readers cannot verify a specific sentence, and AI assistants have nothing to quote.
Average article length1,822 characters5,140 charactersContent answering a real question is nearly three times longer because the answer requires conditions, exceptions, and numbers. The brevity of the first period was not conciseness, but an absence of substance.

Timeline distribution

The gold segment of the bar shows articles from that month answering a real query. The August spike has two drivers: an improved pipeline and twenty legacy articles rewritten manually at the same URLs.

  • April 202615/79 19%
  • May 202620/61 33%
  • June 20267/30 23%
  • July 20266/29 21%
  • August 202617/25 68%

Takeaways from the archive

5
  1. The issue is not that a model writes. The issue is the prompt: "write about something new".When asked for industry novelty, the model fabricates novelty: it invents a name for an unnamed phenomenon and builds an article around it. The result is grammatically correct and completely unsearchable, because no one searches for a term coined yesterday. Given an actual customer question, the same model answers that question. The difference lies in a single paragraph of instructions.
  2. The volume of worthless content grows faster than the return on publishing it.At peak, 83% of the URLs in the sitemap were news articles. The offer pages we wanted search engines to see drowned among two hundred pieces on fabricated phenomena. After cleanup, that share dropped by half, with the exact same number of offer pages.
  3. Deleting is the worst possible move.A URL that was once public should continue to resolve; otherwise, inbound links break. Content without potential receives "noindex, follow", drops out of the sitemap and RSS feed, but stays on the site and within internal linking. The crawler still passes through it to reach offer pages.
  4. Rewriting works better than publishing another piece alongside it.We rewrote the twenty most promising articles under the same URLs: new headline, new body, list of sources. Adding a new article next to an old one increases the number of low-value pages; replacing content under an existing URL reduces it while leveraging a URL already indexed by search engines.
  5. A mechanical gate beats good intentions.The rule "write about things people search for" was in place from day one and never worked. What worked was a rule in the code that evaluates the headline before publication and has no moods. An editorial resolution without enforcement is just a resolution.

How we calculated this

"Answers a real query" sounds subjective and would be if evaluated by a human. Here, an automated rule assesses it, the same one that decides whether content on this page gets indexed. The rule is simple enough to replicate on your own archive:

  • A title passes if it starts with a question word (how much, does, how, when, why, who, where, what) or contains a question mark.
  • It also passes if it includes an amount with a unit, a term from the legal register (GDPR, KSeF, act, contract, consent), or an indicator of an instruction (guide, step-by-step, template, checklist).
  • It fails immediately if it contains an uppercase acronym in parentheses. That is the hallmark pattern of a model-invented name.
  • It fails if the count of English industry terms and mid-sentence capitalized foreign words reaches two or more.
  • The first word of each sentence and all-caps words are skipped; otherwise, every title would trigger a false alarm.

The rule is strict by design and rejects some useful content. That is intentional: the cost of a false rejection is one unindexed article; the cost of a false acceptance is another low-value page in the sitemap. The figure we report is therefore a conservative lower estimate.

What to do on your end

The sequence we followed and would repeat:

  1. Calculate the share. Take twenty titles from the last quarter and search each as an exact phrase. If no one searches for it, the article answers nothing.
  2. Check your sitemap. When more than eighty percent of your URLs belong to the news section, your offer pages drown.
  3. Do not delete. Apply noindex, follow, remove the post from the sitemap and the RSS feed, but keep it on the website.
  4. Rewrite the best pieces at the same URLs.
  5. Change the generator prompt from "write about something new" to a specific customer question. It is one paragraph and costs nothing.

See how we calculate the cost of cleanups and implementations under AI implementation pricing. Our publishing guidelines for machine-generated material are on the AI-generated content page.

Frequently asked questions

Does Google penalize AI-written content?

Not for authorship alone. Google documentation covers generative content separately and does not use the tool as a criterion. The criterion is the page's purpose. Mass-producing pages with no value for readers is penalized, whether written by a model or a human. In our archive, the problem was not a penalty, but silence: 224 pieces and zero traffic, because no one was searching for what we wrote about.

How many of your AI articles answered a real query?

Out of 193 pieces generated up to August 18, 2026, 34 met the criterion: 18%. Out of 11 generated after the pipeline fix: all 11. The criterion is mechanical: the title must start with an interrogative word, include a question mark, an amount, a reference to a regulation, or a tutorial promise, and must not include an invented buzzword.

What exactly did you change in the generator?

Three things. A pre-writing research step searching for an unanswered question instead of a trending topic. A gate blocking any article lacking a list of sources. And a topic queue with known-demand questions prioritized before automated search. The publishing frequency stayed the same: one piece per day.

What should be done with zero-value articles?

Do not delete them. A URL once made public should continue to resolve, otherwise inbound external links are lost. A practical sequence: apply "noindex, follow" to zero-potential content, remove it from the sitemap and RSS feed, but keep it on the site and within internal links. Rewrite the most promising pieces at the same URL. We rewrote twenty.

Where do the numbers in this report come from?

From a single file: the hexart.pl article database as of August 29, 2026. The sample covers 224 public pieces published between April 3 and August 27, 2026. There are no third-party study figures or forecasts here: the report covers one website and discusses only that site.

Citations and contact

Data from this report can be cited without asking for our permission. Please link to this page so readers can verify the methodology. If you run a similar archive and want to compare results, or need the raw dataset, write to us: biuro@hexart.pl.

No-obligation call

Schedule a workshop with Paul

Thirty minutes or a full workshop, choose what suits you. Book a slot directly in the calendar and get instant confirmation.

Paul Lazniak

HEXART Founder