AI News, Explained: Anthropic’s Destroyed Books, OpenAI’s Ten Maths Results and Five More Stories
Two headlines from last week sounded as if they had been written to make everyone angry or breathless.
Anthropic destroyed millions of books. OpenAI solved ten unsolved maths problems for about $2,000.
Both headlines contain a real fact. Neither is a useful account of what happened.
That is the point of this new weekly format. We will take two or three AI stories apart, trace the strongest claims to their evidence, and keep the rest of the week brief. The aim is not to remove the interesting bit. It is to find out which bit is actually interesting.
Anthropic really did destroy millions of books. It did not happen last week
The newest version of the story began with a report about ISBNdb marketing a bulk book-sourcing service to AI companies. The service promised physical books published before the recent flood of AI-written material. It was quickly connected to Anthropic’s much larger book-scanning operation, Project Panama.
ISBNdb later said the proposed service had only tested demand and that it had not bought, scanned or sold a book. Anthropic’s operation was real, but it was mainly carried out from 2024 and had already appeared in a June 2025 court ruling.
What Project Panama actually did
The court’s order in Bartz v Anthropic records two very different routes into Anthropic’s book library.
First came unauthorised digital copies. Anthropic downloaded 196,640 books from Books3 in early 2021, at least five million from LibGen later that year, and at least two million from PiLiMi in 2022.
Then, in 2024, it pursued a route based on purchased print books. Anthropic spent what the court described as “many millions of dollars” buying millions of books, often used. Contractors removed the bindings, cut the pages, scanned them and discarded the paper. One physical book became one searchable PDF in Anthropic’s private research library.
So the physical destruction is not an invented detail. Millions of books were cut apart and discarded after scanning. What is not supported by the evidence is the more lurid claim that Anthropic sought rare or unique books to stop anyone else reading them. The documented purpose was an industrially efficient print-to-digital conversion. The cited court record does not show that Anthropic targeted last surviving copies or bought books to prevent others reading them.
That does not make the operation trivial. It raises legitimate questions about cultural stewardship, secrecy, creator compensation and what happens when a private company buys at a scale that can disturb the second-hand market. Those questions do not need a fictional motive attached to them.
Why destroying the originals mattered in court
Judge William Alsup treated acquisition, library storage and model training as separate uses.
On the record before the court, using the books to train Anthropic’s models was transformative fair use. Converting each purchased print copy into a searchable digital replacement was also fair use, for a different reason: Anthropic retained one copy in a new format, did not keep the physical original as an additional copy and did not distribute the scan.
The judge reached the opposite conclusion about keeping pirate-sourced books in a permanent, general-purpose library. A potentially fair training use did not excuse the earlier acquisition and continuing storage of those files. Buying a physical copy later did not erase liability for the pirated one.
This was a US district-court ruling on a particular factual record. It was not an appellate decision declaring every kind of AI training lawful, and it did not create a general rule that a company must destroy a book before it can scan it.
The settlement created a second, digital destruction story
On 20 July 2026, a court gave final approval to Anthropic’s $1.5 billion copyright settlement. It covers 482,460 eligible works and implies roughly $3,000 per work before fees and costs.
That settlement concerns the alleged piracy of LibGen and PiLiMi files, not the disposal of purchased print books. Its dataset-destruction clause requires Anthropic to delete the original files downloaded from those libraries and copies derived from them within 30 days of final judgment, subject to legal preservation duties. It explicitly excludes scans made from purchased physical books.
As of 3 August, that deadline had not expired, and the public record did not establish whether the deletion had already happened.
The accurate version is still uncomfortable: Anthropic acquired more than seven million pirate-library files, later bought and destructively scanned millions of physical books, and built a private searchable library. But “Anthropic just destroyed rare books so nobody else could read them” combines old events, new legal news and an unsupported motive into one very shareable sentence.
OpenAI published ten serious maths results, not ten identical solved problems
On 1 August, OpenAI published ten advances in mathematics and theoretical computer science. The work spans high-dimensional geometry, coding theory, group theory, operator algebras, quantum complexity, cryptography and extremal combinatorics.
OpenAI says an internal version of Astra, its “next major model”, generated the mathematical arguments. Humans then used the model to prepare the manuscripts, after which the model formalised each argument in the Lean proof assistant. OpenAI released the paper, narrated reasoning traces and Lean files for all ten results.
That is more substantial than a model producing a persuasive-looking answer in a chat window. It is also different from the headline version in four important ways.
| Headline shorthand | What the published evidence supports |
|---|---|
| “Ten unsolved problems solved” | The set mixes resolutions and counterexamples with improved bounds, hardness results and other advances. “Ten new research results” is safer. |
| “The AI worked alone” | OpenAI says Astra produced the mathematical arguments, but people selected the problems and prepared the manuscripts with the model. |
| “All ten are verified discoveries” | The Lean derivations are checkable. Novelty, significance, prior art and the match between each formal statement and the historical problem still require specialist review. |
| “Ten breakthroughs cost $2,000” | OpenAI priced the successful solution-finding tokens at roughly $2,000 using Sol API rates. That excludes model development, human work, validation and unsuccessful work outside the selected set. |
Some of the claims are unusually significant
The collection includes a construction intended to establish the existence of a non-sofic group, a counterexample to Connes’s rigidity conjecture, and new upper bounds in high-dimensional sphere packing. Other results improve bounds or hardness results rather than closing a famous question outright.
The breadth may matter as much as any one result. A single system reportedly moved across research areas that would normally involve different specialist communities. If the work survives scrutiny, that is evidence of a system doing more than retrieving or explaining known material.
It is evidence from one company about an unreleased model, though. The papers appeared only two days before this article. Some chapters credit mathematicians for careful readings and comments, but there has not yet been time for comprehensive independent review.
Lean provides strong receipts, with a boundary
Lean checks whether a formal conclusion follows from stated definitions and assumptions. OpenAI’s formalisation manifest reports zero unfinished sorry placeholders, only Lean’s standard axioms, and a review status of “agent-reviewed”. The repository includes instructions for rebuilding the proofs independently.
That is a meaningful form of verification. It makes a hidden algebraic mistake or skipped logical step harder to smuggle through polished prose.
Lean does not determine whether the encoded theorem perfectly captures the informal historical problem. Nor can it establish that a result is novel, important or free of an overlooked assumption outside the formal statement. Those remain jobs for mathematicians and the normal research process.
The honest verdict is stronger than “another benchmark” and weaker than “AI finished mathematics”. OpenAI has published credible early evidence that a general AI system can produce substantial, formally checkable research across several fields. The proof artefacts justify attention. The age of the work justifies patience.
The EU’s AI transparency rule now applies, but it does not label every AI-assisted post
Article 50 of the EU AI Act began applying on 2 August. The broad claim that every AI-assisted article or social post now needs a conspicuous label is wrong.
The European Commission’s guidance separates the duties of AI providers from the duties of organisations using their systems:
- Providers must ordinarily tell people when they are interacting directly with AI, unless it is obvious, and make synthetic text, audio, image and video detectable with machine-readable marks.
- Deployers must clearly disclose deepfakes and AI-generated public-interest text published without substantive human review or editorial control. They also have duties around emotion recognition and biometric categorisation.
- Standard editing is excluded from the provider marking rule. Public-interest text that receives genuine human review or editorial control does not need the deployer label. A spell-check alone is not substantive review.
- Systems placed on the market before 2 August have a limited grace period until 2 December 2026, but only for the machine-marking and detection requirement.
For a content team, the practical job is not to put “made with AI” on everything. It is to know which tools generated or materially altered an output, retain provenance where possible, and make a real person responsible for checking the substance before publication.
Four more stories worth knowing
LinkedIn and Snap are penalising synthetic finished products, not every use of AI
Snap said wholly AI-generated videos would no longer be eligible for Spotlight recommendations. That is a distribution rule, not an upload ban. Videos enhanced or edited with Snapchat’s AI tools remain eligible and receive transparency indicators.
LinkedIn’s product chief described new and improved classifiers, member reporting and a planned analytics signal for posts readers perceive as inauthentic or heavily AI-assisted. LinkedIn is also replacing “enhance your post” with proofreading intended to preserve the writer’s voice. Some of these are tests or work in progress, not platform-wide features already delivered.
The shared line is useful: both platforms are trying to distinguish AI assistance from output audiences perceive as low-quality or inauthentic. How LinkedIn avoids false positives remains an open question.
Similarweb estimates AI Overviews now appear in more than 40% of US searches
Similarweb’s 2026 Generative AI Landscape report estimates that more than 40% of US Google searches trigger an AI Overview. It also reports that Google searches are 5.4% longer since AI Mode and that an AI mention is associated with a site visit being 2.5 times as likely.
These are third-party measurements, not Google’s internal telemetry, and the last figure does not prove that the mention caused the visit. The direction still matters for content teams. A click is no longer the only useful outcome from search. Citations, brand mentions and whether a source can be retrieved cleanly are becoming part of visibility.
ChatGPT use is crossing job boundaries before job titles change
OpenAI analysed more than 800,000 work-related messages from US ChatGPT users. It classified 16.8% of all work messages as concerning tasks associated with another occupation. Once generic work such as summarising and scheduling was excluded, the figure rose to 43.5% of occupation-specific messages.
The sample suggests AI may change who performs a task before it removes or creates a job. Marketers troubleshoot websites; salespeople inspect datasets; small-business owners draft copy or review contracts.
It is a vendor-produced study of prompts, not a representative labour survey. It does not show whether the work was completed, correct, used or reviewed by a specialist. Treat it as an early signal about task allocation, not a count of jobs replaced.
Google Earth’s one-day image generator rollback was a trust-boundary failure
Google added Nano Banana 2 image generation to Google Earth on 30 July, then rolled it back the following day after people shared generated screenshots that appeared to violate its policies.
The generated images were watermarked and did not change the shared Google Earth view seen by other users. That distinction disappeared when a screenshot travelled without the product interface around it.
The lesson is not that Google secretly replaced satellite imagery. It is that a fictional layer inside a trusted factual product needs guardrails that survive export. The context around an AI output is part of the safety system, until somebody crops it out.
The useful habit is to keep the boundary attached to the claim
Last week’s loudest stories became misleading when one boundary disappeared: purchased books versus pirate files, formal proof versus independent review, provider duties versus publisher duties, AI assistance versus synthetic output, prompts versus completed work, and a private generated layer versus a shared map.
Those distinctions are not pedantry. Each one changes the decision a reader, creator or business should make.
That is what this series will keep doing: identify what changed, show the strongest evidence for it, state what the evidence does not establish, and stop before a useful result turns into mythology.