High-Risk AI Deadlines: When Does Article 26 Apply?

If you run an application classified as a high-risk AI system under the European Union regulatory framework, the operational logging obligations under Article 26(6) will require substantial changes to your backend architecture. A common misconception circulating across industry summaries is that high-risk deployer enforcement took full effect on 2 August 2026. That timeline is legally inaccurate.

Regulation (EU) 2026/1744 (the Digital Omnibus on AI, published in the Official Journal on 24 July 2026 and in force since 27 July 2026) amended Article 113 of Regulation (EU) 2024/1689, extending the timelines for high-risk AI systems. Its recitals state plainly that "the date of application of Sections 1, 2 and 3 of Chapter III is set to 2 December 2027" for systems classified as high-risk pursuant to Article 6(2) and Annex III. Under the amended Article 113, Chapter III Sections 1, 2, and 3, which contain the core technical requirements for high-risk systems and the obligations for providers and deployers in Articles 12, 19, and 26, apply on a deferred schedule:

  • 2 December 2027: Chapter III Sections 1 to 3 apply to stand-alone high-risk AI systems referenced in Article 6(2) and Annex III, such as recruitment scoring, credit evaluation, and critical infrastructure management.
  • 2 August 2028: Chapter III Sections 1 to 3 apply to product-embedded high-risk AI systems referenced in Article 6(1) and Annex I, which undergo third-party conformity assessments under existing Union harmonization legislation.
  • 2 August 2026: The baseline application date where general-purpose AI model obligations in Chapter V, transparency rules under Article 50, and regulatory governance structures became active.

While 2 December 2027 might appear distant, treating it as an engineering buffer is a serious architectural mistake. Article 26(6) mandates a minimum log retention floor of six months. To establish a compliant audit trail that withstands regulatory scrutiny on day one, your logging infrastructure must capture production traffic months before the application deadline. Furthermore, logging cannot be retroactively injected into historical model invocations; the instrumentation must exist at the time of execution. Article 111(2) also leaves an open interpretive question rather than a settled answer: legacy high-risk systems keep their transitional status only until they undergo significant changes in design, and how that limb interacts with the deferred application dates is not resolved on the face of the amendment. Treat it as a question for counsel, not a planning assumption.

The Three Logging Duties: Articles 12, 19, and 26(6)

To build a compliant logging pipeline, engineering teams must decouple three distinct provisions in Chapter III that are frequently conflated. Article 12 is a design duty on the system, Article 19 a provider retention duty, and Article 26(6) a deployer retention duty, so the record-keeping obligations land on different actors in the AI supply chain. Conflating provider design requirements with deployer operational duties leads to fatal gaps in your compliance architecture.

ArticleTarget EntityRetention Figure in the TextPrimary Objective
Article 12System ProviderNo retention period is stated. Article 12(1) sets a capability duty: high-risk AI systems "shall technically allow for the automatic recording of events (logs) over the lifetime of the system".Ensuring system-level technical capabilities for traceability and event detection.
Article 19System ProviderProviders keep the Article 12(1) logs "to the extent such logs are under their control", for a period appropriate to the intended purpose and of at least six months.Retaining infrastructure logs for post-market monitoring and regulatory audit.
Article 26(6)DeployerDeployers keep the automatically generated logs under their control "for a period appropriate to the intended purpose of the high-risk AI system, of at least six months".Preserving operational context, user inputs, human oversight, and downstream outcomes.

Under Article 12(1), the provider must architect the high-risk AI system so that it technically allows for the automatic recording of events over the lifetime of the system. Article 12(2) specifies that these logging capabilities must enable the recording of events relevant to identifying risks within the meaning of Article 79(1), to facilitating the post-market monitoring referred to in Article 72, and to monitoring the operation of high-risk AI systems referred to in Article 26(5). For the systems in point 1(a) of Annex III, Article 12(3) sets an explicit minimum: recording of the period of each use (start and end date and time), the reference database against which input data has been checked, the input data for which the search led to a match, and the identification of the natural persons involved in the verification of the results.

Article 19 obligates the system provider to store these internal system logs for at least six months, but only "to the extent such logs are under their control." In contrast, Article 26(6) places a direct, independent record-keeping obligation on the deployer. You cannot satisfy Article 26(6) by relying on your upstream model supplier's compliance documentation. The deployer is responsible for preserving the operational trace of how the system was actually used in practice.

"Under Your Control": Where the Logging Duty Lands

The central legal and engineering test sits in the words of Article 26(6) itself: deployers "shall keep the logs automatically generated by that high-risk AI system to the extent such logs are under their control". When an engineering team integrates an upstream foundation model or hosted inference API into an internal workflow, teams often assume the logging obligation belongs to the cloud host. In practice, the boundary resolves in the exact opposite direction.

In modern enterprise architectures, the high-risk AI system is rarely a standalone model weights file. The high-risk system is your application: the ingestion pipeline, the retrieval-augmented generation (RAG) vector lookup, the prompt assembly, the model invocation, the post-processing filter, and the human workflow that acts on the output. Within this architecture, the inference API is merely a downstream execution component.

  • Data under Deployer Control: Raw user inputs, assembled system prompts, user metadata, retrieved context chunks, decoding parameters, received model completions, UI presentation timestamps, human-in-the-loop overrides, and downstream business decisions.
  • Data under Provider Control: Physical GPU cluster telemetry, CUDA kernel execution timings, internal scheduler queue states, and hardware fault registers.

An upstream inference host has no visibility into what prompt was submitted before template injection, which database record was retrieved, or what decision a human reviewer made after viewing the completion. If a market surveillance authority requests the operational logs for a disputed automated decision, an export of raw server metrics from an inference host provides zero evidential value. You cannot request operational decision logs from your API vendor because the vendor never possessed the surrounding context.

Zero Data Retention Inference: Why Vendors Lack Logs

The operational necessity of application-side logging becomes absolute when using privacy-preserving, zero-retention inference services. Modern AI infrastructure providers increasingly adopt zero data retention policies to meet enterprise data sovereignty and security standards. In a true zero-retention architecture, the inference host does not persist prompt payloads or generated outputs to non-volatile disk storage.

A serverless inference endpoint of this kind processes incoming requests entirely in volatile GPU memory, holding token states only for the brief duration of the active generation session, so prompts and completions are never written to persistent databases or secondary cold storage. Where an operator makes that claim, it is the operator's own assertion about its architecture, typically with no third-party attestation and no EU AI Act conformity statement behind it. The same property holds for any zero data retention service: it is a characteristic of the architecture, not a single vendor's quirk.

Because a zero-retention endpoint purges all runtime memory buffers as soon as the HTTP connection closes, there is no provider-side log repository to query retroactively. If a deployer fails to capture the input payload and returned completion at the moment of invocation, that event record is permanently lost. Selecting an EU-hosted or privacy-focused inference vendor does not discharge your Article 26(6) logging duty; rather, it makes deploying a dedicated application-side logging layer strictly mandatory.

The Six-Month Floor vs. The GDPR Ceiling

Article 26(6) requires deployers to keep the logs automatically generated by the high-risk AI system, to the extent those logs are under their control, "for a period appropriate to the intended purpose of the high-risk AI system, of at least six months, unless provided otherwise in applicable Union or national law, in particular in Union law on the protection of personal data". Engineering teams must recognise that six months is a statutory minimum floor, not a default configuration setting.

The actual retention duration must be calculated by balancing multiple statutory obligations and regulatory constraints across your operating domain:

  1. Operational and Post-Market Horizon: Article 26(5) requires deployers to monitor operation on the basis of the instructions for use, to inform the provider in accordance with Article 72, to notify the provider, distributor and market surveillance authority and suspend use where the system may present a risk within the meaning of Article 79(1), and to report serious incidents, with Article 73 applying where the provider cannot be reached. If your system operates on an annual auditing cycle, a six-month log lifecycle destroys the evidence those reviews depend on.
  2. Sector-Specific Regulatory Mandates: Article 26(6) states that deployers that are financial institutions subject to internal governance requirements under Union financial services law shall maintain the logs as part of the documentation kept pursuant to that law, so the applicable sector retention period governs rather than the six-month floor.
  3. GDPR Data Minimisation and Storage Limitation: Article 5(1)(c) and 5(1)(e) of the GDPR act as a strict ceiling. Retaining personal data in raw prompts for longer than necessary to achieve the specific purpose of traceability constitutes an unlawful processing violation.

To resolve the tension between the AI Act retention floor and the GDPR ceiling, engineering teams should implement automated log lifecycle policies. While system metadata, token usage metrics, timestamps, and model identifiers can be retained indefinitely, raw personal data in prompt inputs should be encrypted at rest, isolated with strict role-based access controls, and programmatically purged or pseudonymised once the post-market monitoring and audit windows expire.

What the Deployer's Log Record Must Contain

The EU AI Act intentionally avoids prescribing a rigid, universal database schema for deployer logs, acknowledging that high-risk classification spans diverse technical applications from CV parsing to automated medical diagnostics. However, based on the traceability objectives in Article 12(2) and deployer oversight obligations under Article 26(2), we can establish a production-ready baseline schema for structured event capture.

Field CategorySpecific Payload ElementsEngineering Purpose
Session & TimingUTC timestamp (ISO 8601), session UUID, execution duration (ms)Establishes precise event sequencing and chronological order for audit reconstruction.
Model ProvenanceExact model string (e.g., Qwen3-235B-A22B), engine version, system fingerprintProves the exact computational artifact and parameter configuration that evaluated the request.
Input ContextSystem prompt hash, user input payload, RAG retrieval document IDsCaptures the exact factual context and instructions supplied to the model.
Inference ParametersTemperature, top_p, max_tokens, presence_penalty, seedGuarantees deterministic audit conditions when testing for drift or anomalous behavior.
Output & OversightGenerated text completion, finish_reason, human reviewer ID, final approval statusRecords the model output alongside the human-in-the-loop validation record.

A critical architectural vulnerability arises when applications call dynamic routing abstractions. Meta-routing aliases (a router, simple, complex or reasoning entry point) select the target model automatically per request, so the model that actually produced the output is decided at call time and is not visible in the model string you sent. Logging only the generic routing alias fails Article 26 traceability: the record names the gateway, not the artifact that generated the completion. Your application must capture the specific underlying model identifier the API returns, and you should verify against the live API what the response actually exposes before designing the record around it. Where reconstructability matters, pinning an explicit model string is the safer design.

Capturing Logs in Your Application Architecture

Implementing an Article 26(6) logging pipeline does not require re-architecting your entire model integration. Because modern inference providers adhere to standard OpenAI-compatible REST schemas, you can intercept and persist audit records directly at the application call site using middleware wrappers or standard SDK hooks.

When you point your client at an OpenAI-compatible endpoint such as Lyceum Serverless Inference, mapping your AI infrastructure requirements onto it, your application keeps its existing client logic and only the base URL and model string change. By wrapping the SDK client call, you synchronously capture the full transaction payload before returning the completion to downstream application services:

The code below illustrates a production Python implementation using the standard OpenAI client SDK, instrumenting the exact invocation to generate a structured audit record:

  1. Configure the OpenAI client pointing to the base URL https://api.lyceum.technology/api/v2/external/serverless and supply your authorization token.
  2. Assemble the invocation payload, explicitly setting parameters such as temperature, top_p, and the verified model string qwen3-235b-a22b.
  3. Dispatch the request synchronously, capturing both the returned response.choices[0].message.content and the response.model metadata field.
  4. Emit a structured JSON audit log entry containing the input payload, response content, exact model identifier, timestamp, and the human supervisor ID prior to committing the business transaction.

Once structured logs are generated at the application layer, the next architectural milestone is securing the log repository itself. To ensure these records satisfy formal regulatory inspections under Chapter IX market surveillance rules, deployers must implement tamper-evident storage mechanisms, such as append-only cryptographic hash chains and immutable object storage vaults.