AI This article was created with the help of AI.
Who Owes the Annex IV File and When Do the Rules Apply?
Technical documentation under the EU AI Act is frequently discussed as a general compliance ideal, but in production it is a rigid legal deliverable governed by Article 11(1). If you build machine learning systems for the European market, the burden of authoring the Annex IV technical documentation file falls squarely on the entity classified as the provider of a high-risk AI system. Article 11(1) states that the technical documentation shall be drawn up before that system is placed on the market or put into service and shall be kept up to date, and that it shall contain, at a minimum, the elements set out in Annex IV. Under Article 18 the provider must keep it at the disposal of national competent authorities for 10 years after the system has been placed on the market or put into service.
For engineering teams in smaller organisations, Article 11(1) provides a critical operational relief valve: SMEs, including start-ups, may provide the elements of the technical documentation specified in Annex IV in a simplified manner, and the Commission shall establish a simplified technical documentation form targeted at the needs of small and microenterprises, which notified bodies shall accept for the purposes of the conformity assessment. This documentation is distinct from general-purpose AI (GPAI) model documentation, which is governed separately under Annex XI and Annex XII. If your team merely integrates third-party models via inference APIs without modifying their core purpose, you operate as a deployer under Article 26 rather than a provider. However, Article 25 is the trapdoor: your team assumes provider status, and inherits the Annex IV authoring duty, if you put your name or trademark on an existing high-risk system, make a substantial modification to it, or change the intended purpose of a system so that it becomes high-risk high-risk classification.
The statutory timeline for enforcement requires precision. Regulation (EU) 2024/1689 entered into force on 1 August 2024 under Article 113. Following the adoption of Regulation (EU) 2026/1744 (the Digital Omnibus on AI, published in the Official Journal on 24 July 2026 and in force from 27 July 2026), the application dates for Chapter III high-risk obligations were formally extended. The rules for stand-alone high-risk systems under Annex III now apply from 2 December 2027, and those for high-risk AI embedded in physical products under Annex I from 2 August 2028. The earlier date of 2 August 2026 applied only to Article 50 transparency obligations and the Commission's Article 101 fining powers over GPAI model providers. This timeline gives machine learning teams a structured window to embed technical logging, data governance, and architectural mapping directly into their CI/CD pipelines, but the file itself must still exist before the system is placed on the market. This section reflects the consolidated text as verified on 24 August 2026; it is not legal advice, so involve your counsel and compliance function.
In practice four roles matter. A provider that develops a stand-alone Annex III high-risk system and places it on the market under its own name owes the full Annex IV file, all nine numbered points. An SME or start-up provider owes the same elements but may supply them in the simplified manner permitted by Article 11(1), using the simplified technical documentation form the Commission is mandated to establish, which notified bodies must accept for the purposes of the conformity assessment. A party that rebrands, substantially modifies, or changes the intended purpose of a high-risk system inherits the provider's obligations, Annex IV authorship included. A deployer consuming a pre-trained model inside its original intended purpose owes Article 26 duties, chiefly log keeping and use in line with the instructions for use, rather than an Annex IV file. For every provider role, the retention period is the same: 10 years after the system has been placed on the market or put into service, with the documentation kept at the disposal of national competent authorities.
Defining the System, Purpose, and Hardware Components
Point 1 of Annex IV establishes the descriptive foundation of the technical file: a general description of the AI system, which under Point 1(a) means its intended purpose, the name of the provider, and the version of the system reflecting its relation to previous versions. Rather than an abstract summary, this reads as the concrete operational identity of the thing you shipped, and it serves as the technical baseline against which all downstream risk assessments, validation metrics, and boundary constraints are audited.
Interface Architecture and Hardware Dependencies
Under Annex IV Point 1(b) through Point 1(e), the documentation must explicitly define how the AI system interacts with, or can be used to interact with, hardware or software that is not part of the system itself, the versions of relevant software or firmware and any requirements related to version updates, the forms in which the system is placed on the market or put into service, and the hardware on which it is intended to run. If the model consumes real-time telemetry from external microservices or delivers embeddings to a retrieval pipeline, every integration boundary must be specified alongside required update cadences and dependency locks. Point 1(d) names software packages embedded into hardware, downloads and APIs explicitly, so the description of delivery mechanisms has to say which of those applies.
- System Identification and Versioning: Semantic version strings, Git commit hashes, release tags, and technical links to prior production iterations under Point 1(a).
- Intended Purpose and Boundaries: Exact operational domain, target end-user profiles, explicitly unsupported use cases, and boundary constraints.
- Software and Firmware Stack: Host kernel versions, GPU driver levels, container runtime baselines, and serving engine specifications, together with any requirements related to version updates under Point 1(c).
- Hardware Profile: Target accelerator specifications, minimum accelerator memory capacity, required PCIe or NVLink interconnect bandwidth, and CPU host memory ceilings under Point 1(e).
- Deployment Formats and Deployer Interfaces: API contracts (such as OpenAPI schemas), GUI layouts, embedded binary footprints, and structured deployer instruction manuals required by Point 1(g) and 1(h).
Documenting the Development Process and Training Data
Point 2 of Annex IV represents the core technical payload of the file, demanding granular transparency into how the model was designed, trained, and optimised. Point 2(a) asks for the methods and steps performed for the development of the system, including any recourse to pre-trained systems or third-party tools and how those were used, integrated or modified. Point 2(b) asks for the design specifications: the general logic of the system and of the algorithms, the key design choices with their rationale and assumptions, the main classification choices, what the system is designed to optimise for, the expected output and output quality, and the trade-offs made in the technical solutions adopted to comply with the requirements technical requirements of Chapter III, Section 2. In practice that is where you record why a model family, a precision level such as FP8 versus BF16, or a quantisation scheme was chosen over the alternatives.
Data Provenance, Pipeline Lineage, and Compute Accounting
Annex IV Point 2(d) requires data requirements in the form of datasheets, and Point 2(c) requires an account of the computational resources used to develop, train, test and validate the system. The datasheets must describe the training methodologies and techniques and the training data sets used, with a general description of those sets, information about their provenance, scope and main characteristics, how the data was obtained and selected, labelling procedures, and data cleaning methodologies such as outlier detection. In engineering terms that means writing down your filtering heuristics, any synthetic data generation pipelines, and your labelling protocol, then keeping the record versioned alongside the dataset. Point 2(c) also covers the system architecture, explaining how software components build on or feed into each other, which is where GPU cluster topology, interconnect speeds and total training compute belong.
Validation, Testing, and Cybersecurity Measures
Moving from training lineage to system verification, Annex IV Point 2(e) through Point 2(h) covers the human oversight assessment, pre-determined changes, the validation and testing procedures, and the cybersecurity measures put in place. Point 2(g) is explicit on the evidentiary standard: it requires the metrics used to measure accuracy, robustness and potentially discriminatory impacts, plus test logs and all test reports dated and signed by the responsible persons. That is a document with a name on it, not an automated CI dump, so decide early which engineering lead and which quality owner signs.
Evaluation Metrics, Robustness, and Technical Human Oversight
Under Point 2(g), testing protocols must encompass unseen validation datasets specifically curated to detect statistical bias, performance degradation on edge cases, and adversarial vulnerabilities. Point 2(e) demands a technical assessment of human oversight mechanisms compliant with Article 14, requiring engineers to design interpretable outputs, confidence scoring thresholds, and deterministic fail-safe override switches directly into the serving layer.
- Curate and isolate independent validation datasets that mirror real-world demographic and operational variance without leakage from training splits.
- Execute quantitative benchmarking across accuracy, F1-score, perplexity, calibration error, and demographic parity metrics to identify potential discriminatory bias.
- Conduct stress testing and adversarial fuzzing against prompt injection, out-of-distribution inputs, and upstream service timeouts.
- Implement technical oversight hooks under Article 14, including confidence score exposure, human-in-the-loop confirmation gates, and automated circuit breakers.
- Document cybersecurity posture under Point 2(h), detailing model weight encryption, secure key management, runtime isolation, and API authentication controls.
- Compile evaluation artefacts into dated test logs signed by responsible engineering personnel, cataloguing baseline performance and pre-determined continuous update thresholds.
Monitoring, Risk Management, and Performance Metrics
Sections 3, 4, and 5 of Annex IV bridge internal model evaluation with ongoing risk management. Point 3 requires an exhaustive characterization of system capabilities and performance limitations, including explicitly defined error margins across sub-populations and documented scenarios where the model is expected to fail or hallucinate. This technical specification must also define valid input data ranges, expected token distributions, and payload rate limits to prevent operational misuse.
Point 4 is a single line in the text: a description of the appropriateness of the performance metrics for the specific AI system. In practice that means the engineering team has to justify why the chosen metrics fit the domain and risk profile, rather than relying solely on global accuracy figures: why false negative penalties dominate in medical scoring, or why a precision ceiling was set in automated filtering. Point 5 then requires a detailed description of the risk management system in accordance with Article 9, which is where architectural risks, mitigation strategies, and residual risk acceptance levels enter the technical file.
| Annex IV Section | Core Regulatory Focus | Engineering Deliverable | Risk Alignment |
|---|---|---|---|
| Point 3: System Capabilities & Limits | Operational boundaries, error envelopes, and input data requirements | Input constraint specifications, known failure modes, and sub-population error bounds | Prevents out-of-distribution deployment and unsafe edge-case execution |
| Point 4: Metric Appropriateness | Description of the appropriateness of the performance metrics for the specific system | Domain metric analysis justifying loss functions and multi-objective thresholds | Ensures evaluation targets correlate directly with real-world safety |
| Point 5: Risk Management System | Detailed description of the risk management system in accordance with Article 9 | Article 9 risk matrix, mitigation log, and residual risk assessment file | Maintains continuous alignment with high-risk safety mandates |
Lifecycle Changes, Standards, and Post-Market Plans
The final components of Annex IV address system maintainability, standards alignment, and operational surveillance across the post-deployment lifecycle. Point 6 requires teams to maintain a formal audit log of all material changes, retrained weights, hyperparameter adjustments, and pipeline modifications made over the lifetime of the system. This ensures that as models drift or receive fine-tuning patches, the technical file remains continuously synchronized with the running production binary.
Harmonised Standards, EU Declaration, and Post-Market Monitoring
Under Point 7, the provider must list all harmonised European standards applied in full or in part, such as emerging CEN/CENELEC AI standards published in the Official Journal of the European Union, or provide an extensive technical explanation of the bespoke engineering solutions adopted to meet Chapter III requirements. Point 8 mandates the inclusion of a signed copy of the EU declaration of conformity under Article 47 as part of the overarching conformity assessment dossier.
Point 9 concludes Annex IV by requiring a documented post-market monitoring plan in accordance with Article 72. This plan must define active telemetry collection pipelines, automated data drift detection thresholds, feedback loops for reporting serious incidents, and structured processes for continuous model evaluation under real-world traffic conditions.
- Versioned Change Register (Point 6): Granular release logs detailing dataset updates, weight retraining events, pipeline changes, and rollback points.
- Standards Alignment Matrix (Point 7): Direct cross-references to harmonised European standards or technical descriptions of equivalent mitigation architectures.
- EU Declaration of Conformity (Point 8): Fully executed Article 47 declaration confirming system compliance with Chapter III requirements.
- Post-Market Monitoring Plan (Point 9): Operational telemetry strategy under Article 72, defining drift detection metrics, latency logging, incident escalation trees, and retraining triggers.
Mapping Infrastructure Evidence to Your Annex IV File
When assembling an Annex IV technical file, ML engineering teams must maintain a strict technical boundary between what the infrastructure host supplies and what the team must author itself. Infrastructure providers supply raw compute, networking fabrics, and facility security, but they do not hold an EU AI Act conformity statement for your proprietary application, nor can any cloud vendor legally discharge your Annex IV obligations. Provider-supplied evidence covers physical datacentre specifications and hardware parameters, while your team retains exclusive legal authorship of model logic, training lineage, validation logs, and post-market monitoring plans.
At Lyceum, we provide European infrastructure designed to give engineering teams full architectural visibility and hardware determinism for their technical documentation. Teams deploy across sovereign European data centres located in Paris and Finland (EU/EEA), ensuring that hardware descriptions under Annex IV Point 1(e) map directly to dedicated accelerator configurations. By running dedicated models on On-demand GPU VM instances with raw SSH control or utilizing Dedicated Inference behind private, isolated endpoints, engineers maintain direct custody over runtime environments, driver stacks, and audit logs. Furthermore, our architectural support for volatile memory execution enables teams to implement verified zero data retention workflows without secondary disk persistence.
| Annex IV Requirement | Infrastructure Provider Delivers | ML Engineering Team Authors |
|---|---|---|
| Point 1(e): Hardware Profile | Physical accelerator model names and memory configurations as recorded in the catalogue, interconnect speeds, and facility topologies | System VRAM allocation budgets, multi-GPU topology mapping, and throughput scaling requirements |
| Point 2(c): Compute Accounting | Raw VM runtime logs, cluster node provisioning telemetry, and hardware uptime status | Total training FLOPs accounting, GPU-hours logged per run, and distributed framework configurations |
| Point 2(h): Cybersecurity | Upstream data-centre physical security and hypervisor isolation | Model weight encryption, API token authentication, network ingress rules, and container runtime policies |
| Point 9: Post-Market Monitoring | Raw infrastructure health metrics and network availability telemetry | Model prediction latency logs, semantic drift tracking, output error monitoring, and Article 72 incident response procedures |
Building an Annex IV technical file is fundamentally an exercise in production systems engineering. By integrating data versioning, automated evaluation runs, and reproducible hardware provisioning directly into your continuous deployment stack, your team can construct a resilient compliance dossier that satisfies national supervisory authorities without derailing model development cycles.