Connect with us

NEWS

OpenAI Backs a National Web of AI-Ready Bio Data

OpenAI endorsed three bills to centralize AI-ready biological data, the same day its foundation issued more than $125 million to create more of it.

Published

on

OpenAI on Sept. 15 endorsed three bipartisan bills that would format U.S. biological datasets for AI training and park them in a Department of Energy web. Richard Johnson, the company’s national security risk mitigation lead, said the package would help measure how powerful models affect biosecurity and put tools in defenders’ hands.

The same Tuesday, the OpenAI Foundation began issuing more than $125 million to create and open scientific datasets. The safety pitch and the data buy landed together.

A Federal Web of AI-Ready Biological Data

Johnson said the company’s thinking on biology tracks its work on AI and cybersecurity. Risk is one half of the problem. The other half is getting usable tools to people who can stop an attack.

We need to be worried both about the safety and the risk side, but we also need to be looking at the opportunities for getting the tools into the hands of the defenders.

Richard Johnson, OpenAI national security risk mitigation lead

He also said high-quality data and shared standards are what let labs measure model capabilities against everyone else in the industry. Caitlin Frazer, executive director of the National Security Commission on Emerging Biotechnology, a congressional advisory body that is backing the same bills, put the data itself in national-security language: high-quality, AI-ready biological data, she said, must be treated like a strategic national resource.

That is the turn the wire story underplays. OpenAI is not asking Congress to cap training runs. It is asking Congress to industrialize the biological records those runs would train and test on.

Stanford’s Model Already Wrote 16 Working Viruses

On Aug. 6, 2026, researchers at Stanford University and the Arc Institute, a Palo Alto nonprofit, reported in the journal Science that generative models had designed complete viral genomes that worked in the lab. It was the first time those models had written a full genome that could replicate inside cells.

Brian Hie, an assistant professor of chemical engineering at Stanford with a joint appointment at Arc, led the work with first author Samuel King, a bioengineering graduate student. They used genome language models known as Evo 1 and Evo 2, trained on DNA the way chatbots are trained on text, then prompted them from the tiny bacteriophage ΦX174.

THE PHAGE EXPERIMENT

  • The yield: The team synthesized 302 AI-written designs and got 16 fully functional bacteriophages that infected and killed E. coli.
  • The host: The viruses were built to attack bacteria only and, the authors said, pose no threat to people.
  • The method: The models learned gene order and conserved sequence from genomes across life, then wrote new genomes molecule by molecule.
  • The warning: Johns Hopkins biosecurity specialists said existing DNA-order screens are not built to catch sequences that never existed in nature.

That last gap is why a national warehouse of AI-ready genomes cuts both ways. The same well-labeled corpus that helps a defender score a model also helps a genome model learn what a viable virus looks like. Stanford needed public sequence libraries to get to 16 working phages. The bills would make the next library larger, cleaner, and formatted for training.

What the Three Bills Would Do

The three proposals OpenAI named are the Web of Biological Data Act, the AI-Ready Bio-Data Standards Act, and the SCALE Biology Act. Two already have Senate companions. All three sit with the House Science, Space, and Technology Committee. SCALE is the only one ordered reported.

THE THREE BILLS OPENAI ENDORSED

Bill Number Lead agency Job House status
Web of Biological Data Act H.R. 9307 Energy Department National lab builds a single entry point for AI-ready biodata Introduced June 11, 2026
AI-Ready Bio-Data Standards Act H.R. 7907 NIST Define and enforce AI-ready format for qualified federal biodata Introduced March 12, 2026
SCALE Biology Act H.R. 8981 NIST Stand up a biometrology lab for engineering biology and biorisk measures Ordered reported July 21, 2026

Rep. Matt Van Epps (R-Tenn.) and Rep. Jake Auchincloss (D-Mass.) introduced the Energy bill. Sen. Todd Young (R-Ind.), who chairs the biotechnology commission, filed the Senate companion with Sens. Alex Padilla (D-Calif.), Mike Rounds (R-S.D.), and Andy Kim (D-N.J.). Reps. Ro Khanna (D-Calif.) and Jay Obernolte (R-Calif.) introduced the standards bill; Young and Sen. Ben Ray Luján (D-N.M.) have the Senate twin, S. 4069. Rep. April McClain Delaney (D-Md.) led SCALE with Obernolte, Rep. Zoe Lofgren (D-Calif.), and Rep. James Moylan (R-Guam).

Obernolte’s name is on two of the three OpenAI bills and on the separate FRONTIER Act that the company backed the same day for independent third-party audits of models. Chris Lehane, OpenAI’s chief global affairs officer, told reporters he had met Obernolte’s office and could support that independent-verification piece. An OpenAI spokesperson said the company has not endorsed FRONTIER as a whole.

NIST and Energy Get Custody of the Datasets

H.R. 9307 orders the energy secretary, within 180 days of enactment, to pick a national laboratory through a competition and fund a centralized Web of Biological Data. The lab would host federal datasets or pipe in existing databases, tag them for quality, and apply tiered cybersecurity. Phase I, due in two years, is a pilot with a single login, APIs built with the National Institute of Standards and Technology, and a ban on access from adversarial countries and from countries that will not share data back. Phase II, due in five years, would connect any database that holds federally funded data and store it in “a standard quality and format for the hosted biological data that is appropriate for training artificial intelligence models.”

WHAT THE ENERGY WEB WOULD REQUIRE

  • The money: The bill authorizes $30 million plus $310 million for the first three years, then $80 million in each of the next two, $500 million in all.
  • The board: An 11-member advisory panel, four of them from industry, would oversee the build. Federal advisory-committee rules would not apply.
  • The partners: The lab would work with NSF’s National Artificial Intelligence Research Resource, NIST, and the National Library of Medicine, and could sign cost-sharing deals with industry and philanthropy.
  • The lock: A contractor would audit cybersecurity twice a year, and the text says privacy, consent, and human-subjects rules still apply.

H.R. 7907 gives NIST two years to make federally funded biodata AI-ready, then to test those rules with the National Science Foundation so they do not crush smaller labs. “AI-ready,” as the bill defines the job, means a dataset generated and formatted so it can train models and support AI-and-biotech research. SCALE, whose long title is the Standards and Calibration for American Leadership in Engineering Biology Act, would write a biometrology program into the NIST statute, including measurement work on biorisk, biosafety, and biosecurity. It authorizes $55 million in fiscal 2026, rising to $85 million in fiscal 2030, $348 million across those five years.

Young’s commission has been writing this stack since its April 2025 report. In a June 11 statement on the Energy bill, the commission said the United States still has no coordinated way to manage biological data, while China’s approach is to mine public datasets abroad and close its own. Young called that data a strategic national resource for biodata that should be secured from the Chinese Communist Party. Four of the extra board seats on the Energy web are reserved for industry. Frontier labs that already train on biology would have an obvious interest in how those seats are filled.

Five Pathogen Cases Landed in the Same Window

On Sept. 10, Anthropic published a threat report covering December 2025 through August 2026 and described five cases in which people used Claude in ways that could support biological weapons work. The company said it does not assert that those people intended harm. The dual-use problem is the point: the same grant language can underwrite a vaccine or a weapon.

Three of the cases involved viral modification. In May 2026, Anthropic said, a user tied to a military-affiliated institute asked Claude to draft a grant for gain-of-function work on chikungunya, a mosquito-borne virus, including mutations that could help it spread and dodge immunity, then tests in animals. A second user spent weeks planning mammalian-adaptation experiments on highly pathogenic avian influenza. A third had Claude Opus 5 draft, in about an hour, a full grant on orthopoxviruses, the family that includes smallpox and mpox. Two further cases involved venom peptides and toxin redesign. Several users, Anthropic said, routed around regional blocks and hid the purpose of the work.

Two days later, on Sept. 12, Anthropic chief executive Dario Amodei posted a 3,800-word essay, “We Must Pace the Frontier,” arguing that labs should slow the rate at which they improve model capabilities so safety work can keep up. He listed bioterrorism among the risks and said Anthropic would, on its own, give outside evaluators employee-level access. Sam Altman said he agreed with the evaluator idea. OpenAI’s Tuesday endorsements answer that essay with a different instrument: more data, more NIST yardsticks, and a licensed auditor inside the biggest labs, not a freeze on training.

OpenAI had already moved on the physical half of the threat. In June, Altman signed an open letter with Amodei, Demis Hassabis of Google DeepMind, and Mustafa Suleyman of Microsoft AI asking Congress to require DNA and RNA vendors to screen orders, verify buyers, and keep records. That letter tracks the Biosecurity Modernization and Innovation Act from Sens. Tom Cotton (R-Ark.) and Amy Klobuchar (D-Minn.). OpenAI had previously endorsed that bill and a separate Homestake AI Act. Screening the mail-order gene shop still leaves the models, and the datasets they train on, as the open flank.

The Foundation Issued $125 Million in Data Grants

Hours around the same endorsement, the OpenAI Foundation, the nonprofit parent, launched Public Data for Health, its second life-sciences program after an April Alzheimer’s effort. In a post by Abhishaike Mahajan and Jacob Trefethen, who heads life sciences there, the foundation said it was starting with more than $125 million in grants to nonprofits and universities to create and preserve datasets “made broadly available to researchers.”

“We expect many remaining breakthroughs in preventing and curing disease to come from pairing the intelligence of new models with more observations of the world, in other words, more data,” the foundation said. Early money includes work at the University of North Carolina on personalized cancer-vaccine data, OpenADMET competitions on how small molecules move through the body, and an effort to open common technical documents from failed drug programs. Trefethen has said the value of scientific data will only rise as models get better at reading it.

That is the supply-chain problem, not a mystery about molecular biology. Frontier models are data-hungry in a field where every useful observation still comes from a wet lab. The federal bills would force future taxpayer-funded experiments into an AI-ready mold and pipe them through one Energy Department door. The foundation would pay partners to build more of those observations in public. Together they enlarge the corpus. They also raise the price of admission. A lab that can format data to a NIST spec, pass a licensed auditor, and sit on an Energy Department advisory board will clear the new bar. Smaller shops will spend their grants on compliance.

OpenAI Wants Yardsticks Before Any Freeze

Johnson said the point of the bills is to keep measurement in front of capability, so the industry does not “move so fast that we’re getting ahead of our ability to know how these things could be used and what they could do.” That is a slower claim than Amodei’s, and a more convenient one for a company that is still training. Yardsticks let OpenAI argue it is watching biosecurity while the web of data, and the foundation’s grants, keep filling the tank.

THE WEEK THE DATA ARGUMENT LANDED

  1. March 12, 2026: Khanna and Obernolte introduce the AI-Ready Bio-Data Standards Act; Young and Luján file the Senate twin.
  2. May 21, 2026: McClain Delaney introduces SCALE. The Science Committee orders it reported on July 21.
  3. June 11, 2026: Van Epps and Auchincloss introduce the Web of Biological Data Act, matching Young’s Senate bill.
  4. Aug. 6, 2026: Stanford and Arc report 16 working AI-designed bacteriophages in Science.
  5. Sept. 10, 2026: Anthropic details five biology cases in its misuse report.
  6. Sept. 12, 2026: Amodei publishes “We Must Pace the Frontier.”
  7. Sept. 15, 2026: The OpenAI Foundation announces Public Data for Health, Lehane backs FRONTIER’s evaluator rule, and OpenAI endorses the three biodata bills.

None of the three biodata bills has passed the House. SCALE is the only one ordered reported. The Energy web and the NIST standards clock would start only after enactment, then run two to five years. In that gap, genome models will keep training on whatever public sequence they can reach, and the next 16 viruses will not wait for a national lab login.

Harry is the editor of AKRON SCORE, his own independent title, and numbers are the part of the job he takes most seriously. Ten years of newsroom work, reporter first and later editor, taught him that a wrong figure does more damage than a wrong adjective, so every score, percentage, price and headcount is traced to its origin: the official box score, the audited accounts, the published dataset, the spec sheet, the government release. If a number cannot be sourced it does not appear. He writes for readers in every time zone across ten sections, giving sports and gaming the same care as news, business, technology, science, entertainment, lifestyle, travel and auto. Corrections are made in public, under a policy published on the site: the article is amended, the change is dated and described at the foot of the piece, and the original error is not quietly deleted. Readers who find a figure that does not add up can write to support@akronscore.org and he will check it against the source.

Continue Reading
Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending