A Definition Is Being Written: Five Conditions Before Washington Defines “Super Intelligence”

The White House renamed AI and gave itself 60 days to define it. Here’s what any definition should require a system to show before it scales.

Sima Yazdani, Founder, MagiqCircle LLC · Graduate student, ASU College of Health Solutions

October 2026

On September 29, the White House ordered the executive branch to stop saying “artificial intelligence.” Agencies must now write “Super Intelligence” and “SI,” and are told not to acknowledge the older terms “in any applicable setting.”

An illuminated garden path.
An illuminated garden path

Read Section 3 and the rename turns out to be circular. “Super Intelligence” is defined as the technologies already covered by the statutory definition of artificial intelligence. The same machines, under a grander name.

Section 3(b) is the part worth your attention. Within 60 days, around November 28, the President’s science adviser must deliver proposed legislative language for a federal definition of “Super Intelligence,” including whether it should “modify, expand upon, or otherwise supersede” the existing one.

A definition is being drafted right now, on a clock, in one office.

Renaming a technology does not change what it can do, but it can change what gets examined. By putting a statutory definition in play on a sixty-day clock, the order has opened something useful: a window for informing the public, and for drawing in the expertise of the people who build these systems and the people who will live with them.

This essay proposes five conditions any such definition should carry. It is not an argument against building. It is an argument for building things whose claims can be checked, and for widening the circle of people who can check them.

Three words that do not mean the same thing

Artificial intelligence is already ordinary and already defined in law: a machine-based system that makes predictions, recommendations, or decisions influencing real or virtual environments. A scan reader, a translator, a scheduling tool. Narrow, useful, and wrong outside its job.

Artificial general intelligence is a proposal, not a product. OpenAI’s charter calls it systems that “outperform humans at most economically valuable work.” DeepMind’s levels framework measures breadth against depth and has not marked human range as reached. Other frameworks reject the same system. Until the test, the tester, and the appeal are named, “this is AGI” is a claim, not a finding.

Superintelligence is the research term. Nick Bostrom defines it as an intellect that greatly exceeds human cognitive performance in virtually all domains of interest. No deployed system meets that description.

The bars are stacked, not interchangeable. The space in “Super Intelligence” is not the distinction. Style guides disagree about spacing, and English does not let a space change a claim. The distinction is which claim is being made. A product launch, a leaderboard, and a federal naming order can all be heard as the research term while meaning nothing of the kind.

What the tests actually show

A benchmark is not a definition.

ARC-AGI asks a system to infer a rule from a few grids and apply it to a new one. Part of the test set is withheld so a system cannot have trained on the answers. In May 2025, the best scores on that withheld set ran from 3.0 percent down to 0.9 percent, and the benchmark’s own authors note that anything under 5 percent is not meaningful.

By the end of 2025 the number had moved a long way. Claude Opus 4.5 scored 37.6 percent at $2.20 a task. A refinement built on Gemini 3 Pro reached 54 percent, at $30 a task. The best entry under contest rules managed 24.03 percent for 20 cents.

Two things are true at once. Seven months produced real progress. And every task on that benchmark was solved by at least two ordinary people drawn from the general public, who needed a median of roughly two minutes each.

A partial score on one novel-reasoning test falls well short of human-level range. And a score reported without its cost cannot be compared with anything.

Three further gaps are worth naming, because a benchmark measures none of them.

Embodiment. A system that only reads and writes text has never had to pick up an unfamiliar object without crushing it.

World models. A sense of how things behave (that a dropped cup falls, that a full one spills when tilted) is what lets you plan an action you have never performed. Current systems predict untried actions poorly.

Continuous learning. Most systems stop learning at deployment. A model released in March still knows only March.

None of this proves scale cannot get there. It marks what a score has not shown.

Safety is not alignment

These two words get used as if they were one. They are different tests.

Safety asks whether a system stays inside the harm it was allowed to cause: it does the authorized job, it fails visibly outside that job, and it can be stopped.

Alignment asks whether the objective it pursues is the objective the operator intended, rather than a proxy that scores well and misses the point.

A proxy is a stand-in: a number that is easy to measure, standing in for a goal that is not. Reward the number and a system will pursue the number. That is how a model trained to raise a satisfaction score learns to tell people what they want to hear.

A system can be aligned to a bad order. A safe system pointed at the wrong person is still a weapon. Both tests are required.

Alignment without a halt is a claim. Safety without a published failure is a brochure.

The concrete problems were named in 2016: rewards that can be gamed, oversight that does not scale, behavior that changes outside the test. Current practice answers with human feedback, written constitutions, and red-teaming. Those reduce known failures. They do not settle unknown ones.

A constitution that no court enforces

Constitutional AI is a training method, and its name is easy to overread. Anthropic introduced it in 2022 so a model could be steered by a written list of principles instead of human labels on individual outputs.

It is not a constitution in the legal sense. No public adopted it. No court enforces it.

The value of saying so plainly is the demystification: a company can write the rules it wants its model to follow, and a reader can ask who wrote them.

Anthropic’s January 2026 constitution does rank its four properties (broadly safe, broadly ethical, compliant with company guidelines, genuinely helpful), but states that the ranking is “holistic rather than strict.” Higher priorities dominate without deciding. A published ranking that guides rather than decides beats no ranking. It still does not tell an affected person which way their case came out.

An outside team later broke that constitution into 205 testable rules and ran adversarial scenarios against five models. Violation rates fell from 15 percent to 2 percent across generations. Worth knowing, and worth noting that the auditing agent was itself one of those models, and that its authors caution the method “should not be taken at face value.”

That is measured adherence to a written list. It is not proof the list is the right list.

Three distinctions keep this conversation constructive. A principle list is not a law. A model judging another model is not an independent auditor. And alignment to a constitution is not safety against a bad steward: a rule against mass surveillance does nothing if the operator’s objective is mass surveillance.

What should be built

A health digital twin is a good example of the direction, and of what it costs to do properly.

It is a living model of a person, drawn from records, images, wearables, and consented data, used to rehearse a treatment and to catch a decline early. A 2026 survey found privacy and security discussed in 70 of 75 studies. A thematic review found equity, bias, and proof of benefit still open.

So: the model does not act on the body. The person can see the data and withdraw it. The clinician makes the irreversible decision. An independent reviewer can publish a failure that bears on a life.

Those are procedures, and procedures are silent on what counts as harm. The Universal Declaration of Human Rights is where that content lives: privacy and correspondence in Article 12, the freedom to seek and impart information in Article 19, peaceful assembly in Article 20, an effective remedy before a tribunal in Article 8. And it is the one yardstick here that no builder wrote. It binds no one on its own, which is exactly the point: rights name the harm, and an inspector the operator cannot discipline is what makes the naming count.

That is a reason to accelerate. Calling a present twin safe, before that trace exists, is the error this piece is against.

On the same day as the naming order, frontier-company leaders signed a voluntary accord committing to “four layers of controls and audits”: internal controls, an internal oversight team, external audits, and board review. Asked whether it was binding, Trump said he thought it “morally binding.” It is a step toward a standard. It is not yet one.

A finding the affected person cannot read is not accountability.

Who may not audit themselves

Nuclear technology is the precedent this question already has. It is dual-use. It was governed by treaty and inspection rather than by a capability threshold. And its test is structural: a state that declares material submits to verification by an agency it does not control.

The test does not ask whether intentions are good, and it does not wait for a weapon. It asks whether an outside party can look.

The Iranian record shows what happens when the answer is no, and the regime is not the nation. In August 2002 the National Council of Resistance of Iran publicly identified undeclared facilities at Natanz and Arak, a disclosure independently recorded in the nonproliferation literature. The Agency learned of them from media reports, not from Iran. A visit scheduled for October 2002 was postponed and did not take place until February 2003; inspectors reached the site in March. The Board did not find Iran in non-compliance until September 2005, and did not report the case to the Security Council until February 2006.

That gap is the argument.

The pattern has continued. A June 2026 Board resolution recorded that the Agency could not verify previously declared material, including a large quantity of high-enriched uranium. By September, the Agency had been denied inspections for months. Alongside it runs a human record: at least 2,159 executions in 2025, an 88-day national internet shutdown, and facial recognition and phone data used to find protesters.

Two failures run through all of it, and both have direct analogues in AI.

The first is self-audit. A party that has denied inspection of one dangerous capability cannot be the certifier of another. A safety case, a capability claim, or an AGI declaration issued by such an authority is inspection as theater.

The second is standing. Through the executions, the blackout, and the denied inspections, the people most exposed had no forum in which their account counted. An accountability process that excludes the population it is accounting for is not accountability.

The test is structural, not political, and it binds any government, including ones your own country counts as friends. It also scales down. When aircraft certification work was delegated largely to the manufacturer, the 737 MAX reached service with a flight-control system regulators had not fully examined. A lab, a ministry, or a vendor that answers only to itself is running the same arrangement, smaller.

The same test has an information-layer answer, and that one is closer to home for anyone reading this outside Iran. In October 2025 researchers at the Citizen Lab documented a network of more than 50 fake accounts pushing AI-generated video into Persian-language social media: fabricated news clips attributed to a major broadcaster, deepfaked musicians singing altered lyrics, and a synthetic video of a prison being bombed that newsrooms took for real and republished. Who built it is stated as a likelihood, not a finding. Whether it worked, the researchers answer plainly: it did not produce the mobilization it was reaching for.

What it did produce is the part that belongs here. A diaspora already watched by the state it fled now cannot tell which of the accounts arguing with it are people, and threats arrive from nowhere and resolve to no one. In February 2026 an anonymous message tagged ten well-known diaspora figures and warned that the corpses of many would soon have to be found, and an exile who had spent months denouncing a faction of the opposition was killed in Canada, with two followers of that faction later charged with murder. Charges are allegations, not findings, and no one further up has been shown to have known or directed. The cost I mean is paid long before any of that is settled, by people with no way to locate the source of a threat and no institution to take it to. Should be, a system that generates a video marks it, a platform that carries the video keeps the mark, and a person a finding is about can read it and answer it. That is the fifth condition and the fourth, applied where most people actually meet this technology.

Wisdom, written down

Bostrom’s 1998 definition, written sixteen years before the one quoted above, named three things: scientific creativity, general wisdom, and social skills. Two of those three have at least attracted benchmark attempts. Wisdom has attracted none. There is no held-out test set for knowing which problem is worth solving, no leaderboard for restraint, no cost-per-task figure for the judgment that a capability should not be used at all.

So the field is improving quickly at the part that can be scored, and standing roughly where it started on the part that cannot. That asymmetry, more than raw capability, is the frontier.

A thousand years ago, in the passage that opens the Shahnameh before a single king or battle appears, Ferdowsi put wisdom first:

خرد رهنمای و خرد دلگشای · خرد دست گیرد به هر دو سرای

Wisdom is the guide and wisdom opens the heart; wisdom takes you by the hand in both worlds.

Ferdowsi, Shahnameh, “In Praise of Wisdom,” in the Khaleghi-Motlagh critical edition; translation mine.

The claim there is that wisdom decides what the other capabilities are for, rather than being a pleasant addition to them. A benchmark cannot carry that, and neither can a naming order.

Put that way it can sound like a mood, which is precisely why it is worth writing down. Written down, it becomes conditions a system has to meet before it is trusted with a consequential decision.


The five conditions

  1. It does the job it was authorized to do, it fails visibly outside that job, and it can be stopped.
  2. Someone the builder cannot discipline is permitted to look.
  3. The person affected can see the data that bears on them and withdraw it.
  4. A finding about a person is readable by that person.
  5. What the system asserts about the world can be traced to a source, and the source can be checked by someone other than the builder.

None of the five is new in isolation. The first is the safety test. The second is the nuclear safeguards principle: verification by an agency the declaring party does not control. The third and fourth are what the health example requires, and what four layers of controls and audits do not yet supply. The fifth is the test this piece has been applying to everyone else from its first line: a claim becomes a finding only once the test, the tester, and the appeal are named. The first four govern what a system does; only the fifth governs what it says.

Healthcare has already done this once. Clinical ethics rests on four principles, set out by Beauchamp and Childress and taught to every clinician: respect for autonomy, nonmaleficence, beneficence, and justice. What makes them more than a wall poster is that each has something behind it. An intervention performed without consent is actionable whatever good it did. Harm below the standard of care is actionable. Justice reaches care through federal civil rights law and reaches research through a review board that has to find the selection of subjects equitable before a study may begin.

AI governance has the principles and almost none of the machinery: no licensure for a generalist model, no malpractice standard calibrated to one, no body that must find a deployment equitable before it is switched on. Set the four against the five conditions and most of them land. Autonomy is the third. Nonmaleficence is the first. Explicability, which Floridi and Cowls add as a fifth principle because the original four were written for an agent who could be asked for reasons and held to the answer, is the fourth and fifth together. Beneficence lands nowhere, and should not: a test can show that a harm was avoided, not that a benefit reached the person who needed it.

Justice lands nowhere either, and that one is a hole. Every one of the five conditions is satisfiable one person at a time. Each asks what some particular person may see, read, withdraw, or halt, and a system can answer all five correctly for every person it touches while performing measurably worse for one group than another. That failure shows only in aggregate, and nothing in a per-person test obliges anyone to aggregate.

So the second condition has to carry more than it appears to: the person the builder cannot discipline has to be able to examine, and to publish, how the system performs broken out by the populations it acts on. That is where justice belongs here, inside the condition that was already about outside scrutiny rather than beside it as a sixth. Unequal harm is caught by whoever is allowed to look across cases, and never by what any one person is shown about their own.

Trust is not an input to this. It is what the five conditions produce, and only once they hold: a system that can be stopped, an inspector who cannot be disciplined, a person who can withdraw their data, a finding that person can read, and a claim whose source can be checked. Asserted before any of that, trust is advertising. Extended after it, trust is what the whole arrangement is for.

What is proposed is the set, applied together, as the minimum a system should have to show before it scales. It is a proposal, offered to be argued with. Anyone who thinks one of the five is wrong, or that a sixth is missing, has until the Section 3(b) language is drafted to say so.

An invitation

Directing this capability is not a job for vendors and ministries alone.

Take one system you actually know, in a clinic, a lab, a classroom. Ask what it is, as it is. Ask what it should have to show before it scales: safety, security, liberty, fairness, wellbeing. Ask who tests it, and who may not. Hold the five conditions against it and say which hold, which fail, and what they leave out.

Send your response to [comment link or email address]. Responses received before the Section 3(b) proposal is due, around November 28, 2026, will be compiled and shared.

For once the invitation has a specific door, and a date on it.


Sima Yazdani is the founder of MagiqCircle LLC, which provides AI strategy and governance consulting. Over thirty years at Cisco she pioneered its AI-enabled knowledge graph and ontology foundation and designed an enterprise AI risk framework grounded in the NIST AI Risk Management Framework. She is co-author of six patents, and is a graduate student at Arizona State University’s College of Health Solutions, DBH Management Track, where she serves as Industry Collaboration Officer for the AI Society.

Disclosure: the author consults on AI governance for a commercial health digital twin platform. The standards proposed here are ones she is accountable for helping meet.

A fully referenced version of this essay, with 39 sources in APA format, is available on request.

Leave a Reply