AI can make writing faster. In Noy and Zhang’sScienceexperiment with 453 professionals, ChatGPT reduced average completion time for midlevel writing tasks by 40% and increased evaluator-rated quality by 18%.
But a faster draft isn’t the useful metric for evidence-heavy content. If an editor then has to investigate the statistics, citations, product claims and conclusions inside it, part of that gain disappears.
The first draft isn’t the right unit of efficiency. Defensible finished content is.
The method in this article is simple in principle:
Evidence → Claims → Reasoning → Prose → QA
Stop asking one prompt to research, infer, judge and write at the same time.
Yes, AI helped me write this article. It was produced using the same workflow described below.
A note on the method
This is a proprietary content-development approach designed by Egor Kaleynik from many years of work in content marketing, practical use of LLMs for marketing tasks and a broader analysis of where AI-assisted content production breaks down.
The method isn’t presented as a scientifically validated universal workflow. Some of its components are supported directly by research; others are analytical and editorial controls derived from experience, failure analysis and the evidence reviewed for this article.
Author’s public profile describes more than 15 years of experience as a marketing strategist and content marketer in IT.
AI slop is actually two problems
“AI slop” sounds like one problem. It is at least two.
- fabricated facts;
- nonexistent citations;
- outdated information;
- claims that go beyond their sources;
- correlation quietly rewritten as causation;
- recommendations resting on assumptions nobody checked.
A part of Microsoft’s AI slop article titled “Headed to
Modern models are better at evidence-heavy work when they have access to suitable external sources. They are still capable of producing plausible rubbish.
The 2026 Nature paper “Synthesizing scientific literature with retrieval-augmented language models” tested systems on scientific-literature synthesis. The authors report that GPT-4o fabricated citations in 78–90% of cases in their evaluated citation task, while the retrieval-based OpenScholar system substantially improved both correctness and citation performance.
The second problem is editorial slop.
The information may be correct, yet the article is generic, bloated, repetitive, excessively balanced, vague, or built from suspiciously uniform sections.
These two problems need different fixes.
A clever rewrite can’t rescue an invented statistic.
Treat both as “make it sound less AI” and you end up polishing confident nonsense.
The normal workflow creates verification debt
A common AI-writing process looks like this:
- Give the model a topic.
- Ask for an outline.
- Generate a draft.
- Fact-check the draft.
- Rewrite the robotic parts.
- Publish.
It feels efficient because visible production starts immediately.
The hidden problem sits inside step three.
The model may have to choose facts, remember or retrieve them, infer relationships, choose a position, resolve gaps and turn the whole thing into fluent prose at once.
Then you receive sentences.
You can’t tell from their tone whether each sentence is:
- directly supported by evidence;
- derived from known inputs;
- a defensible inference;
- a hypothesis;
- an editorial judgment;
- or unsupported.
Verification therefore runs backwards.
Instead of validating evidence and then allowing a claim into the draft, an editor receives a finished claim and has to reconstruct where it came from.
Anti-crisis measures by Chicago Sun-Times after publishing a list of 15 books to read, 10 of which were non-existent
That’s verification debt.
The issue matters because AI productivity gains are real but task-dependent.
Noy and Zhang found large gains on bounded professional writing tasks: 40% lower completion time and 18% higher evaluated quality. Their experiment doesn’t establish the same end-to-end gain for heavily researched articles requiring
A much larger field experiment, “Shifting Work Patterns with Generative AI,” covered 66 firms and 7,137 knowledge workers. Among treated workers who actually used the tool, the researchers found roughly two fewer hours per week spent on email in the second half of the six-month experiment. They didn’t detect broad changes in the overall quantity or composition of work.
So both statements can be true:
AI can make some writing work substantially faster.
A draft can create enough verification work to reduce that advantage on evidence-heavy tasks.
We don’t have good enough evidence to claim that the method in this article is universally faster end to end.
The hidden variable is verification burden.
Stop asking AI to research, reason and write at once
A more defensible workflow separates four operations:
This looks slower at the beginning because prose appears later.
The goal is to stop expensive mistakes before they acquire paragraphs, transitions, headings and three rounds of editing.
1. Define the job before generating content
Start with the reader problem, not the article title.
“AI content quality” is a topic.
“How can a B2B editor use AI without turning every finished draft into a forensic investigation?” is a job.
- who the content is for;
- what they need to understand, choose, or do;
- what the piece should enable after reading;
- what falls outside scope;
- what happens if an important claim is wrong.
That final question controls the amount of process you need.
A rewrite of supplied copy doesn’t deserve the same evidence machinery as a recommendation involving finance, law, medicine, security, or a client’s reputation.
Process hygiene should scale with risk.
Otherwise rigor becomes bureaucracy with nicer labels.
The distinction also fits current marketing practice. Content Marketing Institute’s 2025 research among 274 technology marketers found widespread generative-AI use, while trust in AI output remained mostly moderate rather than high. The same research reports that differentiation and content quality remain live content-production problems.
Content Marketing Institute: 2025 Technology Content Marketing Benchmarks, Budgets and Trends
2. Establish evidence before asking for factual prose
Once the job is clear, determine which parts of the eventual article depend on external reality.
- prices;
- dates;
- product capabilities;
- statistics;
- quotations;
- regulations;
- scientific results;
- market data;
- named examples.
Collect and inspect those before factual prose is generated.
This is where retrieval-augmented generation (RAG) becomes useful. RAG gives the model external material at generation time instead of relying only on what was encoded during training.
Useful? Indeed but not a truth machine.
The 2025 EMNLP Industry paper “Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards” makes this distinction explicit: even with relevant external context, LLMs can still introduce unsupported information or contradictions.
Read “Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards”
A second 2025 ACL Anthology paper, “Visualizing and Benchmarking LLM Factual Hallucination TendencieseCite dataset and found increased hallucination activity when models were exposed to deceptive or fabricated citations in the evaluated setup
The practical lesson isn’t “retrieval fails” but retrieval is one stage of evidence control, not evidence control itself.
retrieve → inspect → accept or reject → extract support → use
|
Is it primary, secondary, or merely repeating someone else? |
|
|
Does it study the same population, product, period, or situation I am writing about? |
|
|
Are several “sources” actually copies of one original claim? |
|
|
Which passage, table, or result supports the wording I want to use? |
Giving the model ten links doesn’t create ten units of truth.
3. Build a claim ledger
Before drafting, make consequential claims visible.
The ledger doesn’t need to become project-management theater. Five columns can be enough.
- Supported fact: directly established by evidence.
- Derivation: follows deterministically from known inputs.
- Defensible inference: supported by evidence but not directly stated by a source.
- Hypothesis: plausible and testable, but not established.
- Editorial judgment: a reasoned choice about importance, framing, or recommendation.
- Unknown: the available evidence doesn’t justify a stronger conclusion.
Here is what that looks like in practice:
|
ChatGPT made professional writing 40% faster. |
Noy & Zhang studied specific midlevel writing tasks |
“Average completion time fell 40% in the studied writing tasks.” |
||
|
ChatGPT reduced average completion time by 40% in the Noy/Zhang experiment. |
Use with population/task qualification |
|||
|
AI makes B2B article production 40% faster. |
||||
|
Late verification can reduce some of the benefit of fast drafting. |
Productivity evidence + verification mechanism |
“Can reduce,” not “eliminates” |
||
|
Research-first AI content is always faster overall. |
||||
|
Separating evidence from prose makes provenance easier to audit. |
Follows from explicit evidence-to-claim mapping |
Present as a workflow property, not experimental result |
That table does more work than “be accurate” ever will.
Unsupported factual claims don’t earn prose.
They get researched, weakened, labeled uncertain, or removed.
This also prevents citation decoration.
A paragraph isn’t sourced because one citation appears at the bottom. The cited material needs to support the actual wording: scope, date, population, comparison and causal strength.
A citation is evidence, not seasoning.
4. Stress-test the reasoning
A fully sourced article can still be wrong.
This is where many fact-checking workflows stop too early.
Suppose the evidence establishes:
- companies using Process X grew faster;
- those companies also invested more in training.
A model can smoothly conclude that Process X caused the growth.
Every factual sentence can have a real citation while the central causal conclusion remains unsupported.
So before drafting recommendations, expose the reasoning.
For important conclusions, ask:
What else could explain this?
When would this conclusion fail?
What would make the opposite recommendation correct?
Which assumption carries most of the argument?
What changes when the main constraint changes?
This matters most in comparison, recommendation, technical and strategy content.
If the article says “choose A over B,” identify the situation where B wins.
If you can’t find one, you may have discovered a powerful universal law.
More often, you have forgotten a variable.
5. Then let AI draft
After evidence, claims and major analytical decisions exist, drafting becomes narrower.
That’s a feature.
“Write the definitive article about X.”
The effective instruction becomes:
“Here is what is supported. Here is what is inferred. Here is what remains uncertain. Here is the decision logic. Explain it clearly without inventing additional factual conclusions.”
The model gets less freedom, excellent.
Freedom is useful during ideation. It is less charming when applied to statistics.
The model can still organize, explain, compare, compress, vary sentence structure, improve transitions, as well as translate technical material into readable language.
What it shouldn’t casually do is expand the factual universe of the article.
If the evidence establishes five things, the draft shouldn’t mysteriously know seven.
6. Run a claims pass and a prose pass
Don’t mix them.
Ignore whether the article sounds good.
- statistics;
- dates;
- quotations;
- named entities;
- product capabilities;
- causal language;
- absolutes;
- comparisons;
- recommendations;
- citations.
Use the same taxonomy as the ledger.
For each consequential statement, decide whether it is:
- supported fact;
- derivation;
- defensible inference;
- hypothesis;
- editorial judgment;
- unknown or unsupported.
Self-critique can still be useful for finding possible defects.
It isn’t independent verification.
Stechly, Valmeekam and Kambhampati’s ICLR 2025 paper, “On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks,” compared GPT-4 self-critique with verification from a sound external reasoner across three reasoning/planning domains. In those experiments, self-critique produced significant performance collapse, while sound external verification produced substantial gains. The scope is specific, so this shouldn’t be generalized into “self-critique never works.”
Use AI self-critique to generate suspects.
Use sources, calculations, tools, tests, or qualified humans to decide whether those suspects are guilty.
Don’t appoint the same model judge, jury, witness and source.
Now stop fact-checking and inspect the writing itself. Look for generic introductions, repeated definitions, obvious implications, mini-summaries after every section, artificial balance, meaningless modifiers, paragraphs that simply restate their heading, ceremonial conclusions, repetitive transitions, identical section patterns and vague abstractions.
Pay particular attention to phrases such as “improve efficiency,” “enhance engagement,” or “offer greater flexibility.” They sound useful because they point toward a benefit without explaining one. Ask what actually became faster, what observable behavior changed, who gained flexibility and under which constraint. Good editing replaces those abstractions with mechanisms, observations, decisions, or actions.
This is how AI slop looked a couple of years ago
A sentence doesn’t need to be measurable in a literal numerical sense, but it should earn its place. If the reader can’t observe what it means, picture the mechanism, act on it, or follow the logic behind it, inspect it again.
A before-and-after example
Here is a small demonstration using one of the sources already cited in this article.
Raw generation
“Research proves AI makes professional content production 40% faster while improving quality by 18%.”
It sounds clean but it’s also too broad.
Claim extraction
The sentence contains at least three claims:
- research established a 40% speed improvement;
- research established an 18% quality improvement;
- those results apply to “professional content production” generally.
The already-mentioned Noy and Zhang experiment does support the first two numbers.
It doen’t test professional content production as a general workflow. It studied 453 college-educated professionals completing specific occupation-related writing tasks.
The third claim therefore fails the scope check.
“AI makes professional content production 40% faster.”
The cited experiment doesn’t establish that.
Inference correction
The experiment shows that AI can produce substantial speed and quality gains on some bounded professional writing tasks.
“Can” matters.
“So does every researched B2B article” would be our invention.
Constrained redraft
In Noy and Zhang’s experiment, professionals using ChatGPT completed the studied writing tasks 40% faster on average, while evaluator-rated quality increased by 18%. The result demonstrates substantial gains on those tasks, not a universal 40% productivity gain for content production.
A little stiff, accurate and traceable.
Prose edit
AI can make writing much faster. In Noy and Zhang’s experiment, professionals finished the assigned writing tasks 40% faster on average, while quality scores rose 18%. That doesn’t mean your next research-heavy article will be 40% faster. The study measured bounded writing tasks, not the entire content-production process.
Bam, the same evidence but better prose!
That’s the whole method in miniature:
generate candidate → extract claims → test sources → correct inference → constrain draft → edit prose
The AI-avoidant writer doesn’t have to trust the model. But the process is designed so trust isn’t the control mechanism.
You don’t need the full process every time
A rigorous workflow becomes sloppy in its own way when people apply it mechanically.
Use a lightweight process when:
- facts are already supplied;
- the output is low-risk;
- errors are easy to spot;
- the task is transformation rather than discovery.
Add explicit evidence gathering when the content depends on current or external facts.
Add a claim ledger when statistics, citations, product capabilities, causal claims, or consequential recommendations matter.
Add reasoning tests when the article compares, predicts, diagnoses, or recommends.
Move AI away from primary drafting when the value comes mainly from:
- firsthand expertise;
- interviews;
- distinctive judgment;
- confidential knowledge;
- high-consequence interpretation.
Sometimes the best AI-writing workflow isn’t letting AI do much writing.
That doesn’t mean anti-AI, more like choosing the tool according to the job instead of choosing the tool first and inventing a justification afterward.
There is good reason for quality-sensitive professionals to be cautious. In a 2025 Trint survey inaccurate AI output was cited as a challenge by 75% of the surveyed journalists, ahead of reputational risk at 55%. This is journalism evidence, not direct evidence about content marketers, but it shows how quickly verification and reputation become operational issues in accuracy-sensitive publishing
Digiday: “Journalists are using generative AI tools without company oversight, study finds”
Audience trust creates another constraint. Reuters Institute research reports continuing public scepticism about heavily AI-produced news and expectations that AI may make news less accurate, transparent and trustworthy. Again, news isn’t B2B content marketing. The relevance is reputational: readers don’t automatically interpret AI involvement as neutral.
The real distinction isn’t AI-written versus human-written
Humans can write unsupported rubbish perfectly well. AI doesn’t own that market. The more useful distinction is controlled versus uncontrolled content production.
AI changes the economics because fluent text is cheap. That’s useful, but it also makes plausible connective tissue, generic explanations, untested examples and confident conclusions cheap. The problem isn’t merely that an LLM can be wrong. It can make an unresolved question look finished, turning uncertainty into prose before anyone has established whether the underlying claim deserves to survive.
A controlled workflow prevents that flattening. Evidence stays evidence; inference stays inference; hypotheses remain hypotheses; and editorial judgments don’t dress themselves up as empirical facts. When the available sources don’t justify a stronger conclusion, unknown remains an acceptable answer.
That last option matters more than most prompt engineering.
The six-stage version
If you want the method without building a small Ministry of Content Operations, use this:
1. Define the job
Specify the reader, decision, outcome, boundaries and cost of being wrong.
2. Establish the evidence
Find and validate the external information required for factual claims.
3. Build the claims
Separate supported facts from derivations, inferences, hypotheses, editorial judgments and unknowns.
4. Test the reasoning
Look for alternative explanations, negative cases, trade-offs and conditions that reverse the conclusion.
5. Draft from constrained material
Use AI to express approved evidence and reasoning without casually adding new factual conclusions.
6. Run two QA passes
Check epistemic quality first. Then, check editorial quality separately.
The bottleneck was never simply typing. AI has become very good at producing sentences. The harder problem is deciding which sentences deserve to exist.
Sources and further reading
- Noy, Shakked and Whitney Zhang. “Experimental evidence on the productivity effects of generative artificial intelligence.” Science, 2023. 453 professionals; average completion time decreased 40% and evaluated quality increased 18% on the studied midlevel writing tasks. Science article
- Asai et al. “Synthesizing scientific literature with retrieval-augmented language models.” Nature, 2026. Scientific-literature synthesis, citation hallucination and retrieval-grounded generation. Nature article
- Dillon, Jaffe, Immorlica and Stanton. “Shifting Work Patterns with Generative AI.” NBER Working Paper 33795. Field experiment across 66 firms and 7,137 knowledge workers. NBER PDF
- Tamber et al. “Benchmarking LLM Faithfulness in RAG with Evolving Leaderboards.” EMNLP Industry Track, 2025. Evidence that unsupported information and contradictions remain possible in retrieval-augmented generation. ACL Anthology paper
- Mao et al. “Visualizing and Benchmarking LLM Factual Hallucination Tendencies via Internal State Analysis and Clustering.” IJCNLP/AACL, 2025. Introduces FalseCite for studying hallucination behavior around misleading or fabricated citations. ACL Anthology paper
- Guaglione, Sara. “Journalists are using generative AI tools without company oversight, study finds.” Digiday, 2025. Reports survey findings on inaccurate output, reputational risk, privacy and newsroom AI use. Digiday article
- Content Marketing Institute. “2025 Technology Content Marketing Benchmarks, Budgets and Trends.” Technology-marketer data covering GenAI adoption, trust, perceived output quality, workflows and content-production challenges. Content Marketing Institute research
- Reuters Institute. “Generative AI and News Report 2025.” Research on public use of and attitudes toward generative AI, including trust and expectations around AI-supported news. Reuters Institute report PDF
- Stechly, Valmeekam and Kambhampati. “On the Self-Verification Limitations of Large Language Models on Reasoning and Planning Tasks.” ICLR 2025. Compares model self-critique with sound external verification in reasoning and planning tasks. OpenReview paper PDF
