EU AI Act: first systemic-risk GPAI evaluations due September 15
The week of September 10–17, 2026 sits inside the first real enforcement cycle of the EU AI Act's GPAI regime. Sequence of events (FACT unless labelled): 31 July 2026 (announcement, pre-window context): Commission press release "Commission starts enforcing AI Act rules and new transparency requirements" — as of 2 August 2026 the Commission (via the AI Office) can exercise its GPAI enforcement powers: request information and documentation (Article 91), conduct evaluations of GPAI models (Article 92), request measures including risk-mitigation and market-recall steps (Article 93), and impose fines of up to 3% of global annual turnover or €15 million, whichever is higher (Article 101). Algorithmic-transparency rules (Article 50: chatbots, deepfakes) became enforceable on the same date. 1 September 2026: the Commission sent its first information requests to 30+ AI companies, in two streams: (1) safety- and security-related requests under Article 91 (systemic-risk models: evaluations, risk mitigation, incident reporting, cybersecurity) and (2) copyright/transparency requests under Article 53 (training-content summaries, copyright compliance policy). (COMMISSION ACTION per law-firm and policy-blog reporting — Allegiance Law, Magica, EU Perspectives; the requests themselves were announced by the Commission, i.e., a company/institutional claim, echoed by secondary sources.) 8 September 2026: TechTimes reports OpenAI filed the first serious-incident report under Article 55(1)(c) with the AI Office, with OpenAI's chief scientist reportedly admitting a monitoring gap in incident detection. (SECONDARY REPORTING per TechTimes; not independently confirmed with the AI Office. INTERPRETATION: the first public exercise of the GPAI serious-incident-reporting pipeline.) 10 September 2026: POLITICO Europe covers the EU's response to frontier-AI "extinction warnings", including the AI Office's contracting of evaluation capacity with METR (Model Evaluation and Threat Research) — the office is building the technical muscle to use its new Article 92 evaluation power credibly. 15 September 2026 (the in-window event): per Enterprise DNA (8 Sep 2026) and the September press wave, providers of systemic-risk GPAI models (presumed over the ~10^25 FLOP training-compute threshold) were due to submit their first-round systemic-risk model evaluations to the AI Office — red-teaming/adversarial-testing methodology and results (Art 55(1)(a)), systemic-risk assessments (Art 55(1)(b)), plus the related transparency package: energy-consumption disclosures and copyright training-summary template compliance (Art 53). On the same date the first compliance inspection wave of high-risk AI systems began across the Union — targeting HR/resume screening, retail banking/credit scoring, and healthcare/AI triage — co-led by the French CNIL, German BfDI and Spanish AESIA together with the AI Office, spanning the 24 Member State market surveillance authorities (MSAs). 16 September 2026: Lawfare (Eliška Andrš) publishes "You Don't Have to Sell It to Be Bound by It — GPAI and the EU AI Act", a governance analysis arguing Chapter V binds providers of systemic-risk GPAI models even where the model is deployed internally rather than "placed on the market" — significant because GPAI obligations are typically framed around market placement, yet Articles 51/55 attach systemic-risk duties to the model's high-impact capabilities regardless of how it is made available.

Tailored emphasis while keeping the full article available.
⌘ Jump to architecture, developer details, and the hands-on route.
The essential information in 30 seconds
The week of September 10–17, 2026 sits inside the first real enforcement cycle of the EU AI Act's GPAI regime. Sequence of events (FACT unless labelled):
- 31 July 2026 (announcement, pre-window context): Commission press release "Commission starts enforcing AI Act rules and new transparency requirements" — as of 2 August 2026 the Commission (via the AI Office) can exercise its GPAI enforcement powers: request information and documentation (Article 91), conduct evaluations of GPAI models (Article 92), request measures including risk-mitigation and market-recall steps (Article 93), and impose fines of up to 3% of global annual turnover or €15 million, whichever is higher (Article 101). Algorithmic-transparency rules (Article 50: chatbots, deepfakes) became enforceable on the same date.
- 1 September 2026: the Commission sent its first information requests to 30+ AI companies, in two streams: (1) safety- and security-related requests under Article 91 (systemic-risk models: evaluations, risk mitigation, incident reporting, cybersecurity) and (2) copyright/transparency requests under Article 53 (training-content summaries, copyright compliance policy). (COMMISSION ACTION per law-firm and policy-blog reporting — Allegiance Law, Magica, EU Perspectives; the requests themselves were announced by the Commission, i.e., a company/institutional claim, echoed by secondary sources.)
- 8 September 2026: TechTimes reports OpenAI filed the first serious-incident report under Article 55(1)(c) with the AI Office, with OpenAI's chief scientist reportedly admitting a monitoring gap in incident detection. (SECONDARY REPORTING per TechTimes; not independently confirmed with the AI Office. INTERPRETATION: the first public exercise of the GPAI serious-incident-reporting pipeline.)
- 10 September 2026: POLITICO Europe covers the EU's response to frontier-AI "extinction warnings", including the AI Office's contracting of evaluation capacity with METR (Model Evaluation and Threat Research) — the office is building the technical muscle to use its new Article 92 evaluation power credibly.
- 15 September 2026 (the in-window event): per Enterprise DNA (8 Sep 2026) and the September press wave, providers of systemic-risk GPAI models (presumed over the ~10^25 FLOP training-compute threshold) were due to submit their first-round systemic-risk model evaluations to the AI Office — red-teaming/adversarial-testing methodology and results (Art 55(1)(a)), systemic-risk assessments (Art 55(1)(b)), plus the related transparency package: energy-consumption disclosures and copyright training-summary template compliance (Art 53). On the same date the first compliance inspection wave of high-risk AI systems began across the Union — targeting HR/resume screening, retail banking/credit scoring, and healthcare/AI triage — co-led by the French CNIL, German BfDI and Spanish AESIA together with the AI Office, spanning the 24 Member State market surveillance authorities (MSAs).
- 16 September 2026: Lawfare (Eliška Andrš) publishes "You Don't Have to Sell It to Be Bound by It — GPAI and the EU AI Act", a governance analysis arguing Chapter V binds providers of systemic-risk GPAI models even where the model is deployed internally rather than "placed on the market" — significant because GPAI obligations are typically framed around market placement, yet Articles 51/55 attach systemic-risk duties to the model's high-impact capabilities regardless of how it is made available.
The Act-vs-implemented deadline distinction (FACT + INTERPRETATION): Article 113(b) made Chapter V (Articles 51–56) applicable from 2 August 2025 — so model evaluations, adversarial testing, serious-incident reporting and cybersecurity were already legal obligations before September 2026. What changed in the current window is that the enforcement machinery became operative (2 August 2026 powers), the Commission began issuing Article 91 requests (1 September 2026), and the 15 September milestone is the first concrete administrative submission date for evaluation documentation inside that machinery. No calendar date for evaluation submissions appears in the Act itself; "September 15" is the AI Office's implemented first-milestone date as reported by the press. Any statement that "the AI Act requires evaluations by September 15" as a statutory deadline is therefore an oversimplification of a real operational deadline.
- The first enforceable frontier-model compliance deadline in the world (discovery's why_it_matters): California's SB 53/AB 54-style laws and voluntary commitments set precedents, but no other jurisdiction had, by Sep 2026, combined a hard model-evaluation obligation with information powers, an inspection apparatus and a turnover-linked fine. The EU is defining what "regulated frontier AI" means operationally.
- Evaluation evidence is now a regulated asset: adversarial-testing outputs move from safety-team internal documents into regulator-submitted records with legal consequences — reshaping how labs resource red-teaming.
- A working model for supervision: the AI Office–METR contract (POLITICO Europe) shows the EU buying third-party evaluation capacity, effectively outsourcing the "ARIA-style" technical scrutiny the Act's authors knew the office could not build in-house — a model other regulators will copy.
- Copyright enforcement enters the frontier: Article 53 training-content-summary requests put EU copyright law into the frontier-training pipeline — the second front of this enforcement cycle alongside safety.
- Signal to enterprises and deployers: inspection wave = deployer-side diligence becomes enforceable now, in the highest-visibility sectors (HR, banking, healthcare) — not in 2028.
CONFIRMED
- Story ID: S34
- Title: EU AI Act: first systemic-risk GPAI evaluations due September 15
- Organization: European Commission (AI Office / Directorate-General for Communications Networks, Content and Technology, digital-strategy)
- Category: governance
- Event date: 2026-09-15 — CONFIRMED in-window (2026-09-15 ∈ [2026-09-10, 2026-09-17])
- Evidence status: CONFIRMED (event and legal framework corroborated by official Commission pages and press releases, the primary Regulation (EU) 2024/1689 text via EUR-Lex and the AI Act Explorer, plus multiple independent outlets — Lawfare, POLITICO Europe, Enterprise DNA, Slaughter and May, CSA)
- Confidence: High
- What it is: The first operational milestone of the EU AI Act's enforcement of obligations on general-purpose AI (GPAI) models with systemic risk. Providers whose models are presumed to present systemic risk (training compute above the ~10^25 FLOP threshold, Article 52) were due by September 15, 2026 to have their first-round systemic-risk model evaluations — including adversarial testing documentation under Article 55(1)(a) — in the AI Office's hands, as part of the first enforcement wave that became legally operative on 2 August 2026, running in parallel with the first coordinated compliance inspections of high-risk AI systems in HR, retail banking and healthcare across 24 Member States. Critical nuance (FACT vs INTERPRETATION): the Act itself contains no calendar date for evaluations — Article 55 obligations have applied since 2 August 2025 (Article 113(b)); September 15, 2026 is the implemented/operational first submission milestone inside the AI Office's supervision agenda, as reported by Enterprise DNA and mirrored across the September press (adversarial-testing/red-teaming methodology reports, energy-consumption disclosures, copyright training-summary template compliance).
What happened?
The week of September 10–17, 2026 sits inside the first real enforcement cycle of the EU AI Act's GPAI regime. Sequence of events (FACT unless labelled):
- 31 July 2026 (announcement, pre-window context): Commission press release "Commission starts enforcing AI Act rules and new transparency requirements" — as of 2 August 2026 the Commission (via the AI Office) can exercise its GPAI enforcement powers: request information and documentation (Article 91), conduct evaluations of GPAI models (Article 92), request measures including risk-mitigation and market-recall steps (Article 93), and impose fines of up to 3% of global annual turnover or €15 million, whichever is higher (Article 101). Algorithmic-transparency rules (Article 50: chatbots, deepfakes) became enforceable on the same date.
- 1 September 2026: the Commission sent its first information requests to 30+ AI companies, in two streams: (1) safety- and security-related requests under Article 91 (systemic-risk models: evaluations, risk mitigation, incident reporting, cybersecurity) and (2) copyright/transparency requests under Article 53 (training-content summaries, copyright compliance policy). (COMMISSION ACTION per law-firm and policy-blog reporting — Allegiance Law, Magica, EU Perspectives; the requests themselves were announced by the Commission, i.e., a company/institutional claim, echoed by secondary sources.)
- 8 September 2026: TechTimes reports OpenAI filed the first serious-incident report under Article 55(1)(c) with the AI Office, with OpenAI's chief scientist reportedly admitting a monitoring gap in incident detection. (SECONDARY REPORTING per TechTimes; not independently confirmed with the AI Office. INTERPRETATION: the first public exercise of the GPAI serious-incident-reporting pipeline.)
- 10 September 2026: POLITICO Europe covers the EU's response to frontier-AI "extinction warnings", including the AI Office's contracting of evaluation capacity with METR (Model Evaluation and Threat Research) — the office is building the technical muscle to use its new Article 92 evaluation power credibly.
- 15 September 2026 (the in-window event): per Enterprise DNA (8 Sep 2026) and the September press wave, providers of systemic-risk GPAI models (presumed over the ~10^25 FLOP training-compute threshold) were due to submit their first-round systemic-risk model evaluations to the AI Office — red-teaming/adversarial-testing methodology and results (Art 55(1)(a)), systemic-risk assessments (Art 55(1)(b)), plus the related transparency package: energy-consumption disclosures and copyright training-summary template compliance (Art 53). On the same date the first compliance inspection wave of high-risk AI systems began across the Union — targeting HR/resume screening, retail banking/credit scoring, and healthcare/AI triage — co-led by the French CNIL, German BfDI and Spanish AESIA together with the AI Office, spanning the 24 Member State market surveillance authorities (MSAs).
- 16 September 2026: Lawfare (Eliška Andrš) publishes "You Don't Have to Sell It to Be Bound by It — GPAI and the EU AI Act", a governance analysis arguing Chapter V binds providers of systemic-risk GPAI models even where the model is deployed internally rather than "placed on the market" — significant because GPAI obligations are typically framed around market placement, yet Articles 51/55 attach systemic-risk duties to the model's high-impact capabilities regardless of how it is made available.
The Act-vs-implemented deadline distinction (FACT + INTERPRETATION): Article 113(b) made Chapter V (Articles 51–56) applicable from 2 August 2025 — so model evaluations, adversarial testing, serious-incident reporting and cybersecurity were already legal obligations before September 2026. What changed in the current window is that the enforcement machinery became operative (2 August 2026 powers), the Commission began issuing Article 91 requests (1 September 2026), and the 15 September milestone is the first concrete administrative submission date for evaluation documentation inside that machinery. No calendar date for evaluation submissions appears in the Act itself; "September 15" is the AI Office's implemented first-milestone date as reported by the press. Any statement that "the AI Act requires evaluations by September 15" as a statutory deadline is therefore an oversimplification of a real operational deadline.
What changed?
- First enforceable frontier-model compliance milestone in the world (discovery's core claim): earlier deadlines (prohibitions Feb 2025; GPAI obligations Aug 2025) had no operational supervision behind them; September 2026 is the first time frontier labs face a concrete, inspector-backed submission date for evaluation evidence, backed by a functional fining power (Art 101).
- Hands-on supervision of frontier labs became real: the AI Office now runs an actual enforcement cycle — information requests (Art 91), evaluations (Art 92), requested measures (Art 93) — against the operators of the most capable models, with METR-style evaluation capacity contracted in (POLITICO Europe).
- Adversarial testing moved from research practice to regulated deliverable: red-teaming results are now documentation a lab must produce on a schedule, to a standard the regulator can interrogate — not just a safety-benchmark scorecard.
- Inspection wave for high-risk AI systems began in parallel: HR, retail banking and healthcare deployers (not just model providers) now face MSA inspections requiring evidence of AI Act compliance documentation — the first sectoral enforcement of the deployer-side regime (postponed Annex III obligations to 2 Dec 2027 do not stop inspections of systems whose obligations applied from Aug 2025/2026).
- Copyright/transparency enforcement started: first Article 53 information requests (training-content summaries, copyright policy) pulled frontier providers into the copyright pipeline the Act created — a separate fight from safety evaluation.
Before → Change → After
- Before: Chapter V rights and duties existed on paper since 2 Aug 2025; the GPAI Code of Practice (assessed adequate; signatories incl. OpenAI, Anthropic, Google, Meta, Microsoft) was the voluntary vehicle; the AI Office could recommend but not compel; no information requests, no inspections, no fines for GPAI providers (Art 101 excluded from the 2 Aug 2025 penalties date and became operable with the general application from 2 Aug 2026); serious-incident reports were a contractual, not enforceable, expectation.
- Change (2 Aug–15 Sep 2026): enforcement powers live (2 Aug); first Article 91 information requests to 30+ companies (1 Sep); first Article 55(1)(c) serious-incident report filed (8 Sep); systemic-risk evaluation submissions due (15 Sep); first MSA inspection wave in HR/banking/healthcare begins (15 Sep); evaluation capacity being contracted (METR).
- After: frontier labs operate under a standing EU supervision cycle — submissions, evaluations, follow-up requests, and the credible prospect of Article 101 fines (€15M/3% worldwide turnover) for intentional/negligent failure; evaluation evidence becomes a regulated input into both EU enforcement and the labs' own public safety narratives; deployments of high-risk AI in EU HR, banking and healthcare face inspector scrutiny; the precedent exists for the rest of the world's regulators (US, UK, Canada, Japan) to benchmark against.
How it works
⌘ For BuilderLegal mechanics (FACT, per Regulation (EU) 2024/1689 as amended by the Digital Omnibus, Regulation (EU) 2026/1744):
- Classification (Arts 51–52): a GPAI model is classified as presenting systemic risk either by Commission decision or — presumptively — where cumulative training compute exceeds 10^25 FLOP (Art 52; criteria in Annex XIII; Scientific Panel can alert the Commission under Art 90). Designation is by implementing act; providers may rebut with documented justifications (Art 51(2)).
- Obligations (Art 55(1)(a)–(d)): providers of systemic-risk GPAI models must, in addition to Arts 53–54: (a) perform model evaluation in accordance with state-of-the-art standardised protocols and tools, including conducting and documenting adversarial testing, to identify and mitigate systemic risks; (b) assess and mitigate possible systemic risks at Union level, including their sources, from development, placing on the market, or use; (c) track, document and report serious incidents and corrective measures without undue delay to the AI Office (and national authorities as appropriate); (d) ensure adequate cybersecurity for the model and its physical infrastructure.
- Enforcement powers (Arts 88–94): the AI Office supervises GPAI providers (Art 88); monitoring actions (Art 89); power to request documentation and information (Art 91); power to conduct evaluations of the model, including with independent experts / third-party evaluation support such as METR (Art 92); power to request measures, including risk mitigation and market recall (Art 93); procedural rights of providers (Art 94).
- Fines (Art 101): up to €15 million or 3% of total worldwide annual turnover, whichever is higher, when the Commission finds intentional or negligent infringement of GPAI provisions, failure to comply with Art 91 requests or supply of incorrect/incomplete/misleading information, failure to comply with Art 93 measures, or failure to make the model available for an Art 92 evaluation. Court of Justice retains unlimited jurisdiction (Art 101(5)). Per the Commission's own framing and Praxikon's reading, Article 101 became operable on 2 August 2026 (excluded from the earlier Chapter XII application date).
- Compliance routes (Arts 55(2), 56): adherence to the GPAI Code of Practice assessed as adequate demonstrates compliance (not a full presumption of conformity) until harmonised standards are published; non-signatories must show alternative adequate means (e.g., a gap analysis vs the Code).
- Timing (Art 113; Omnibus amendments): in force 1 Aug 2024; prohibitions from 2 Feb 2025; Chapter V GPAI from 2 Aug 2025; most remaining provisions (incl. GPAI enforcement powers) from 2 Aug 2026; new Omnibus-added Article 5 prohibitions (sexual imagery/CSAM-related) from 2 Dec 2026; Annex III high-risk obligations deferred to 2 Dec 2027 and Annex I to 2 Aug 2028 by the Omnibus — Chapter V GPAI untouched by the deferral; models on the market before 2 Aug 2025 must fully comply with Chapter V by 2 August 2027 (Art 111(3)).
- The Sep 15 milestone: reported by Enterprise DNA (8 Sep 2026) and the September press wave as the first submission deadline for systemic-risk evaluation documentation (adversarial-testing methodology + results under Art 55(1)(a), systemic-risk assessment under 55(1)(b), plus energy-consumption and copyright training-summary transparency under Art 53) — an implementation date set inside the AI Office's supervision agenda rather than a date found in the Regulation (INTERPRETATION; the source basis is press reporting, not a directly verified AI Office notice).
Why it matters
- The first enforceable frontier-model compliance deadline in the world (discovery's why_it_matters): California's SB 53/AB 54-style laws and voluntary commitments set precedents, but no other jurisdiction had, by Sep 2026, combined a hard model-evaluation obligation with information powers, an inspection apparatus and a turnover-linked fine. The EU is defining what "regulated frontier AI" means operationally.
- Evaluation evidence is now a regulated asset: adversarial-testing outputs move from safety-team internal documents into regulator-submitted records with legal consequences — reshaping how labs resource red-teaming.
- A working model for supervision: the AI Office–METR contract (POLITICO Europe) shows the EU buying third-party evaluation capacity, effectively outsourcing the "ARIA-style" technical scrutiny the Act's authors knew the office could not build in-house — a model other regulators will copy.
- Copyright enforcement enters the frontier: Article 53 training-content-summary requests put EU copyright law into the frontier-training pipeline — the second front of this enforcement cycle alongside safety.
- Signal to enterprises and deployers: inspection wave = deployer-side diligence becomes enforceable now, in the highest-visibility sectors (HR, banking, healthcare) — not in 2028.
What became possible?
- The AI Office can compel evaluation documentation, run its own model evaluations (Art 92) with third-party support (METR), and order risk-mitigation measures (Art 93) — a full supervisory loop over the most capable models.
- Serious-incident transparency: Article 55(1)(c) reports create a regulator-facing record of frontier-model incidents (first exercised by OpenAI on 8 Sep 2026 per TechTimes).
- Turnover-proportional fines make compliance economically non-optional for the largest labs.
- Downstream deployers get an authoritative compliance signal to anchor procurement and risk due diligence (who has submitted, what gaps the AI Office flags).
- Non-EU regulators gain a live experiment on which to calibrate their own regimes (US federal preemption fights, UK's pro-innovation posture, Japan's soft-law approach).
Implications
⌘ For BuilderTechnical
- Evaluation methodology is now a legal artifact: "standardised protocols and tools reflecting the state of the art" (Art 55(1)(a)) — the reference point is the Code of Practice's safety-and-security chapter; expect harmonised standards (EU) to crystallise shortly, and expect the AI Office to push for shared evaluation protocols (METR-style) across labs for comparability.
- Red-teaming capacity gap exposed: OpenAI's own reported monitoring gap (TechTimes, 8 Sep) underscores that even top labs lack systematic incident detection — the gap the AI Office's Art 89 monitoring and Art 92 evaluations are designed to probe.
- Energy disclosure becomes compliance data: energy-consumption reporting (Art 53-related transparency) is now a supporting compliance deliverable — environmental claims and compute-scale claims will be cross-checked against FLOP designations.
- FLOP-threshold disputes will surface: as enforcement bites, labs will litigate the 10^25 FLOP presumption and the definition of training compute boundaries (what counts, clustering of models, hybrid/agent training runs) — technical interpretation becomes legal argument (cf. Cambridge Commentary's Article 51 paper).
- Internal-deployment scope question (Lawfare, 16 Sep): whether "placing on the market" is required at all for systemic-risk duties is contested; if the Lawfare reading prevails, internally-deployed frontier models face the same obligations — widening the technical surface under enforcement.
Developer
- Frontier-lab engineering teams must now operate evaluation pipelines as auditable, dated, versioned deliverables, with reproducible adversarial-testing documentation (protocol, scope, results, mitigations) ready for Art 91 requests — treat the Code of Practice safety chapter as the de-facto spec until harmonised standards land.
- Incident-reporting hooks required: Art 55(1)(c) demands "track, document and report… without undue delay" — labs need automated serious-incident detection/reporting pipelines with AI Office contact paths; OpenAI's 8 Sep filing is the template others will be measured against.
- Open-source nuance: free/open-source GPAI models are exempt from most Chapter V obligations unless they present systemic risk (Art 53/54 recitals) — open-weight systemic models (e.g., a future open-weights >10^25 FLOP release) carry the full obligation set, which will shape open-model release decisions.
- Non-signatories face a documentation gap: providers who did not sign the Code must justify alternative means of compliance (e.g., CoP gap analysis) — the AI Office can and likely will use Art 91 to test those justifications.
- Authorised representatives (Art 54): third-country providers need a mandated EU representative able to produce Annex XI technical documentation on request for 10 years — a live procurement item for non-EU labs.
Enterprise
- Procurement diligence is now anchored in enforcement data: enterprises adopting GPAI-based tools can (and should) verify provider evaluation submissions, incident reports and AI Office measures; that information becomes the due-diligence baseline, replacing vendor self-certification.
- HR, banking and healthcare deployers are in the first inspection wave: EU deployers of AI for resume screening, credit scoring and clinical triage must evidence risk-management, data governance, human-oversight and logging documentation (Arts 8–15) under MSA inspection — the enforcement wave is deployer-facing, not just model-facing; the Omnibus postponement of Annex III obligations to Dec 2027 does not suspend MSAs' powers to inspect systems whose duties already apply.
- Cross-border operations: US and Asian enterprises with EU subsidiaries face dual supervision (AI Office at model layer + national MSAs at deployer layer); central compliance teams need a single evidence repository feeding both.
- Legal-risk budgeting: Art 101 fines (€15M/3% turnover) sit above GDPR-level exposure for the largest labs; enterprises should treat GPAI non-compliance of material vendors as a concentration risk in vendor risk registers.
Strategic
- The EU is now the world's frontier-AI regulator by default: with the US Congress stalled (per the broader week's reporting) and the UK/Japan non-intervening, Brussels is the only jurisdiction actively supervising frontier models in 2026 — the "Brussels effect" applied to AI capability governance.
- Regulatory capture and credibility: the AI Office's dependence on contracted evaluators (METR) and on labs' own red-teaming begs the question of independence; how the office handles its first Art 92 evaluations will set or damage its credibility.
- Divergence risk with the US: the enforcement wave sharpens EU–US divergence (EU: binding evaluations + fines; US: voluntary commitments, state-law patchwork), creating compliance arbitrage questions for global labs and possibly friction in transatlantic AI cooperation.
- The extinction-risk politics arrived inside the machinery (POLITICO Europe): "extinction warnings" are no longer just a public debate — they are the justification for the enforcement architecture and for METR-type capacity, tying existential-risk discourse to concrete regulatory spending and powers.
- Precedent-setting cascades: other regulators (Canada's AIDA, Japan, UK, Korea) will copy or contest the EU's evaluation-first enforcement; the Sep 15 milestone is the reference event they will measure themselves against.
Risks & limitations
- Deadline confusion / legal certainty (INTERPRETATION): press increasingly reports "September 15" as if it were a statutory date; the Act has no such date, and the AI Office has not (as of this research) published an individually-verifiable notice for the milestone. Providers and advisors may over- or under-prepare depending on which reading they follow; a legally challenged first enforcement action on a non-statutory deadline would be an own goal.
- Enforcement without public evidence: no public register yet confirms which providers submitted by Sep 15 or what the AI Office received; reporting (Enterprise DNA, TechTimes, Winzheng) is the only window into compliance status — a transparency gap the office must close to sustain credibility.
- Inspection-wave friction: 24 MSAs with heterogeneous capacity and inclinations (CNIL/BfDI/AESIA lead) create enforcement-arbitrage and inconsistency risk; the Art 81/83 safeguard procedures may be tested early.
- Industry pushback / legal challenge: expect appeals over Art 52 FLOP presumptions, Art 55(1)(a) "state of the art" methodology, and the internal-deployment scope question (Lawfare's analysis flags precisely this battleground).
- Trade-secret vs transparency tension: Art 55 documentation is confidential (Art 78), but the Act's credibility depends on visible enforcement; over-secrecy will look like capture, under-secrecy will trigger litigation.
- Evaluation-capacity dependency (POLITICO Europe): outsourcing core evaluations to METR-type contractors introduces its own risks — contractor availability, methodology disputes, neutrality perceptions — especially at the "extinction-level" stakes the office itself advertises.
- Over-reach/under-reach balance: if the first wave produces no visible consequences, the threats (Art 101 fines) ring hollow; if it escalates to a headline fine quickly, expect accusations of politicised enforcement — the office's first Art 101 decision will be closely judged.
- No directly verified AI Office notice for "September 15": the milestone date is established via press reporting (Enterprise DNA 8 Sep 2026 as the primary written statement, echoed by Winzheng and others); the AI Office's own published calendar was not located during research.
- Incident-report fact rests on one secondary outlet: the OpenAI Article 55(1)(c) filing (TechTimes, 8 Sep) is single-source; the AI Office has not confirmed receipt, and the "chief scientist admits monitoring gap" quote is a secondary paraphrase of a company figure (label: COMPANY CLAIM, reported).
- Compliance status of individual labs unknown: no verified list of which providers were designated/presumed systemic-risk or which submitted by Sep 15; the "30+ companies" receiving RFIs is the Commission's own figure through law-firm summaries.
- Effect of the Digital Omnibus on the wave is partially interpretive: the deferral of Annex III obligations (to 2 Dec 2027) and Annex I (to 2 Aug 2028) is confirmed (Regulation (EU) 2026/1744; Praxikon), but how MSAs sequence inspections of not-yet-fully-applicable Annex III systems is a legal reading, not a settled practice.
- The discovery-named Engineering & Technology (8 Sep) article could not be located/verified and was therefore not used; no URL for it is listed in the sources artifact to avoid fabricating a citation.
- Art 101 applicability date (2 Aug 2026) follows the Commission's framing and Praxikon; one secondary source (RegulatoryAI) gives a different (implausible) date (2 Dec 2027) — the Commission-aligned reading is used and the discrepancy is noted here.
Open questions
- Which providers were (a) presumed systemic-risk by the 10^25 FLOP threshold and (b) formally designated by implementing act — and which actually submitted evaluation documentation by Sep 15, 2026?
- What exactly did the Sep 15 submission package require, and is that requirement documented in an AI Office notice, a Code of Practice annex, or only in press reporting?
- Will the AI Office publish evaluation outcomes or a compliance register, or stay behind Article 78 confidentiality?
- When will the first Article 92 evaluation run, with which external evaluator (METR or another), and on which model?
- Will the first Article 101 fine land in 2026–27, and on which ground (infringement, info-request failure, requested-measure failure, or denial of Art 92 access)?
- Does the internal-deployment reading (Lawfare) survive judicial contact — i.e., are internally-deployed systemic models bound by Chapter V?
- How will FLOP-threshold disputes be adjudicated (training-compute definitions, model clustering, rebuttal evidence)?
- Which sectors get inspected second (after HR/banking/healthcare), and what enforcement findings will the 24 MSAs actually surface?
- How will the Code of Practice be updated (Art 56(8)) now that enforcement is live, and will harmonised standards for GPAI evaluation emerge before the next milestone?
- What will the pre-2-Aug-2025 grandfathered models (full Chapter V compliance due 2 Aug 2027) do differently versus new releases?
What should you do with this?
⌘ For BuilderCircle 1 = providers of systemic-risk and other GPAI models (OpenAI, Anthropic, Google/DeepMind, Meta, Microsoft, xAI, Mistral AI, plus non-EU frontier labs with EU authorised representatives), the AI Office, and the 24 national market surveillance authorities (CNIL, BfDI, AESIA leading).
- Impact: Highest. This is the first enforceable compliance cycle for the most capable models; evaluation documentation, incident reporting and cybersecurity are now regulator-verifiable obligations with turnover-linked fines; the AI Office is directly supervising a handful of firms.
- Recommended action: Frontier labs should treat the Code of Practice safety chapter as the interim spec, stand up versioned, reproducible evaluation-documentation pipelines and automated Art 55(1)(c) incident reporting, and pre-position rebuttal evidence on FLOP thresholds; the AI Office should publish a milestone calendar and a high-level compliance register to convert press-reported dates into verifiable public facts; MSAs should publish inspection-wave criteria and findings consistent with Art 81 safeguard procedures.
Circle 2 = downstream developers and deployers of GPAI-based systems (HR-tech, retail-banking, healthcare, legal/compliance software), evaluation and red-teaming vendors (incl. METR-type organisations), AI Act consultancies and law firms, and EU digital-sector industry associations.
- Impact: High. The inspection wave makes deployer-side compliance evidence an operational requirement in three sectors immediately; evaluation vendors gain a regulated market; advisors gain a differentiated, high-margin service line around Art 55 compliance and CoP gap analyses; downstream developers inherit compliance obligations from the model layer.
- Recommended action: Deployers in HR/banking/healthcare should audit their risk-management/data-governance/human-oversight documentation now and pre-empt MSA inspections; evaluation vendors should standardise on Art 55(1)(a) "state of the art" protocols and seek AI Office recognition paths; consultants should build CoP-gap-analysis and Art 91-request-response playbooks as productised offerings.
Circle 3 = global regulators (US federal/state, UK, Canada, Japan, Korea), insurance and capital markets, civil society and the general public, and the broader AI-industry narrative.
- Impact: Moderate-to-high over 12–36 months. The EU's enforceable evaluation regime becomes the reference experiment for global frontier-AI governance; public trust dynamics shift as serious-incident reporting becomes routinised; investors in frontier labs must price EU enforcement risk (Art 101 fines) into models.
- Recommended action: Non-EU regulators should track the first Art 92 evaluations and Art 101 decisions as calibration data for their own regimes (avoiding copy-paste where EU-specific); insurers should develop GPAI-compliance-linked products; the Commission should invest in plain-language reporting of inspection findings and incident stats to sustain public legitimacy; investors should treat signed Code adherence + submitted evaluations as a new due-diligence checkbox.
- Art 55(1)(a) evaluation-as-a-service: independent adversarial-testing and evaluation-documentation services, aligned with the Code of Practice safety chapter and future harmonised standards — a defensible, regulation-mandated service category with METR's contract validating the model.
- CoP gap-analysis and alternative-means compliance: for non-signatory providers, a concrete, sellable deliverable the AI Office itself references (compare against Codes assessed as adequate).
- Article 91 response / information-request playbooks: high-margin compliance services for the 30+ companies now receiving requests — legal + technical response engineering (Annex XI documentation, copyright summaries, energy disclosures).
- Deployer inspection-readiness for HR/banking/healthcare: audit + remediation offerings for MSAs' first inspection wave — a genuine near-term revenue window for consulting firms (cf. Grant Thornton's positioning).
- Serious-incident monitoring tooling: automated AI-incident detection/reporting pipelines satisfying Art 55(1)(c) "without undue delay" — a product gap OpenAI's own monitoring admission highlights.
- EU authorised-representative services (Art 54): for non-EU labs entering the EU market — a recurring, low-competition compliance business.
VERIFY (executed — see labs/S34.md)
A governed legal event with verifiable statutory specifics (same discipline as S25's statute-check): the research included claim-by-claim verification of the September-2026 enforcement wave against the primary Regulation text (Article 51/52/53/54/55/56, 88–94, 101, 113 of Regulation (EU) 2024/1689 via the AI Act Explorer and the Commission's AI Act Service Desk), the Commission's enforcement FAQ, and the Digital Omnibus (Regulation (EU) 2026/1744) penalty-structure analysis. The lab output is a verification table plus a reusable Article 55(1)(a)–(d) evaluation-submission package outline (consulting artefact). See labs/S34.md for the full table and checklist.
What happens next?
- Late Sep–Oct 2026 (near-term, FACT-adjacent/prediction): AI Office review of Sep 15 submissions; first documented Art 92 evaluations likely in Q4 2026 (with METR or similar); follow-up Article 93 requests where evaluations find unmitigated systemic risks; MSA inspection feedback in HR/banking/healthcare flowing to the Board and Safeguard procedures as needed.
- 2 Dec 2026: new Article 5 prohibited practices (sexual-imagery/CSAM-related, added by the Omnibus) become applicable — the next hard date after the Sep milestone.
- 2026–27 enforcement escalations (PREDICTION): first Article 101 fines before end-2027 (likely on Art 91 request failures or Art 92 access denials first, since those are easiest to prove); one or more legal challenges to FLOP presumptions at the General Court; a visible AI Office clash with a non-signatory provider over "alternative adequate means".
- 2 Aug 2027: full Chapter V compliance for models placed on the market before 2 Aug 2025 (Art 111(3)) — a second, broader compliance cliff hitting older flagship models.
- 2 Dec 2027 / 2 Aug 2028: Annex III and Annex I high-risk obligations apply (post-Omnibus dates); deployer-side enforcement in full swing; harmonised standards for GPAI evaluation likely consolidated by then.
- 2027+ (PREDICTION): Code of Practice updated (Art 56(8)) with enforcement-learned refinements; other jurisdictions mirror or differentiate from the EU's evaluation-first model; the Sep-2026 wave is retrospectively seen as the moment EU frontier-AI enforcement became operational rather than symbolic.
Editorial takeaway
September 15, 2026 is the day the EU AI Act stopped being a compliance exercise and became an enforcement cycle. The headline — "first systemic-risk GPAI evaluations due" — is directionally true and globally significant, but the honest version contains two corrections the reporting mostly elides: the Act itself never sets a September 15 calendar date (the obligations have bound providers since August 2025; the deadline is the AI Office's implemented first-milestone inside a machinery that only became operational on 2 August 2026), and the actual enforcement teeth — information powers, regulator-run evaluations, requested measures, and €15M/3%-of-turnover fines — are all from Article 101-inflected Articles 88–94 that are being exercised for the very first time. Watch three things: whether the AI Office converts its milestone into published, verifiable facts (it has not yet); whether the first Article 92 evaluation and the first Article 101 fine are credible enough to matter; and whether the scope fight — does "placing on the market" even matter for systemic-risk duties, per Lawfare's analysis — lands in Luxembourg. In one sentence: the world's first enforceable frontier-model compliance deadline is real, it is operational, and the next twelve months will decide whether it becomes the global template or a cautionary tale in transatlantic divergence.
Evidence discipline summary: FACT (statutory provisions and dates as enacted — Art 51/52/53/54/55/56, 88–94, 101, 113 of Regulation (EU) 2024/1689; 2 Aug 2025 / 2 Aug 2026 / 2 Aug 2027 / 2 Dec 2026 / 2 Dec 2027 / 2 Aug 2028 dates; Digital Omnibus Regulation (EU) 2026/1744 deferrals; Art 101 fines 3%/€15M; Commission enforcement powers from 2 Aug 2026 per official press release); COMPANY CLAIM (Commission announcement of info requests to 30+ companies, echoed by law-firm summaries; OpenAI chief-scientist monitoring admission as reported by TechTimes); INDEPENDENT EVIDENCE (Lawfare scope analysis; POLITICO Europe on AI Office–METR; Enterprise DNA on the Sep 15 milestone and inspection wave; Slaughter and May and CSA enforcement-readiness briefings; Praxikon penalty-structure analysis); INTERPRETATION (Sep 15 as implemented/operational milestone vs statutory date; internal-deployment scope; Annex III inspection sequencing); PREDICTION (first Art 101 fines, FLOP-threshold litigation, CoP update, transatlantic divergence). The discovery-named E&T article could not be verified and was excluded. No rumours used.
