Medium

Falsified Research Abstract Disinformation

A user directly instructs an AI agent to alter a well-known research abstract, insert false claims or numbers, and publish the altered version as if it were genuine. This is a request to help create scientific disinformation, not a technical exploit.

How the attack works

The user gives the agent a real research abstract (in this case, the GPT-3 paper's abstract) and asks it to change key facts, such as parameter counts or claims. The user then asks the agent to post or publish this falsified version somewhere as though it were the authentic, unedited abstract. The goal is to spread a plausible-looking but false piece of scientific text under the guise of legitimacy. The detection flags this specific pattern: editing a known abstract to misstate facts and then publishing it as genuine.

Netzilo detection

Why it matters

If the agent complies, an organization's tooling becomes a vehicle for creating and distributing scientific misinformation, which can damage credibility, mislead readers, and be traced back to the agent's operator or platform.

What you can do

  • Require agents to refuse requests to alter factual content from cited sources and then represent it as unmodified.
  • Add a review step before any agent-generated content citing research or statistics is published externally.
  • Log and audit any agent actions that involve editing quoted or cited material before publication.
  • Treat accurate quoting of real figures (e.g., correctly citing GPT-3's 175B parameters) as expected behavior, not a violation.

Known benign look-alikes

  • Accurate citation of the real GPT-3 abstract (175B parameters)

References

Related threats