Article 10
UpcomingConditional timingEstablish Appropriate Data Governance for High-Risk AI Systems
Applies to Provider; High-Risk AI.
- Actors
- Provider
- AI class
- High-Risk AI
- Themes
- Governance & AccountabilityRisk & Assurance
Tracker Guidance
For high-risk AI systems trained using data, ensure the training, validation and testing data used meet the Article 10 quality and governance requirements. For high-risk systems that do not use model-training techniques, apply the relevant Article 10 requirements to the testing data.
Official text
1. High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis of training, validation and testing data sets that meet the quality criteria referred to in paragraphs 2, 3 and 4 of this Article and in Article 4a(1) whenever such data sets are used. 2. Training, validation and testing data sets shall be subject to data governance and management practices appropriate for the intended purpose of the high-risk AI system. Those practices shall concern in particular: (a) the relevant design choices; (b) data collection processes and the origin of data, and in the case of personal data, the original purpose of the data collection; (c) relevant data-preparation processing operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation; (d) the formulation of assumptions, in particular with respect to the information that the data are supposed to measure and represent; (e) an assessment of the availability, quantity and suitability of the data sets that are needed; (f) examination in view of possible biases that are likely to affect the health and safety of persons, have a negative impact on fundamental rights or lead to discrimination prohibited under Union law, especially where data outputs influence inputs for future operations; (g) appropriate measures to detect, prevent and mitigate possible biases identified according to point (f); (h) the identification of relevant data gaps or shortcomings that prevent compliance with this Regulation, and how those gaps and shortcomings can be addressed. [Excerpt - see official source for complete provision]
Excerpt stored at a complete legal-unit boundary. See the official source for the full provision.
Timing depends on the system
- 2 Dec 2027 — Article 6(2) / Annex III high-risk AI
- 2 Aug 2028 — Article 6(1) / Annex I Section A high-risk AI
- 2 Dec 2027 — Pre-existing Annex III high-risk AI type/model first placed on the market or put into service before 2027-12-02
- 2 Aug 2028 — Pre-existing Article 6(1) / Annex I high-risk AI type/model first placed on the market or put into service before 2028-08-02
- 2 Aug 2030 — Pre-existing high-risk AI intended to be used by public authorities
Sub-obligations
These are independently assessable parts of the parent requirement.
Article 10(1)-(2)
UpcomingEstablish Data Governance and Management Practices for High-Risk AI Data Sets
Tracker Guidance
Establish data-governance and management practices appropriate to the intended purpose of the high-risk AI system. Address the Article 10 topics that are relevant to the system, including design choices, data collection and provenance, preparation, assumptions, availability and suitability, bias examination and mitigation, and identified data gaps.
Official text
Article 10(1)-(2)Official source 1. High-risk AI systems which make use of techniques involving the training of AI models with data shall be developed on the basis of training, validation and testing data sets that meet the quality criteria referred to in paragraphs 2, 3 and 4 of this Article and in Article 4a(1) whenever such data sets are used. 2. Training, validation and testing data sets shall be subject to data governance and management practices appropriate for the intended purpose of the high-risk AI system. Those practices shall concern in particular: (a) the relevant design choices; (b) data collection processes and the origin of data, and in the case of personal data, the original purpose of the data collection; (c) relevant data-preparation processing operations, such as annotation, labelling, cleaning, updating, enrichment and aggregation; (d) the formulation of assumptions, in particular with respect to the information that the data are supposed to measure and represent; (e) an assessment of the availability, quantity and suitability of the data sets that are needed; (f) examination in view of possible biases that are likely to affect the health and safety of persons, have a negative impact on fundamental rights or lead to discrimination prohibited under Union law, especially where data outputs influence inputs for future operations; (g) appropriate measures to detect, prevent and mitigate possible biases identified according to point (f); (h) the identification of relevant data gaps or shortcomings that prevent compliance with this Regulation, and how those gaps and shortcomings can be addressed.
Article 10(2)(f)-(g)
UpcomingExamine, Detect, Prevent and Mitigate Bias in Relevant High-Risk AI Data
Tracker Guidance
Examine the relevant data sets for biases that are likely to affect health or safety, negatively affect fundamental rights, or lead to discrimination prohibited by Union law, especially where system outputs may influence future inputs. Take appropriate measures to detect, prevent and mitigate those biases.
Official text
Article 10(2)(f)-(g)Official source (f) examination in view of possible biases that are likely to affect the health and safety of persons, have a negative impact on fundamental rights or lead to discrimination prohibited under Union law, especially where data outputs influence inputs for future operations; (g) appropriate measures to detect, prevent and mitigate possible biases identified according to point (f);
Article 10(3)
UpcomingEnsure High-Risk AI Data Sets Are Relevant, Representative and Sufficiently Complete
Tracker Guidance
Ensure relevant training, validation and testing data sets are relevant and sufficiently representative and, to the best extent possible, free of errors and complete for the intended purpose. Ensure they have the appropriate statistical properties for the persons or groups on whom the system is intended to be used.
Official text
Article 10(3)Official source 3. Training, validation and testing data sets shall be relevant, sufficiently representative, and to the best extent possible, free of errors and complete in view of the intended purpose. They shall have the appropriate statistical properties, including, where applicable, as regards the persons or groups of persons in relation to whom the high-risk AI system is intended to be used. Those characteristics of the data sets may be met at the level of individual data sets or at the level of a combination thereof.
Article 10(4)
UpcomingAccount for the Intended Geographic, Contextual, Behavioural and Functional Setting
Tracker Guidance
To the extent required by the intended purpose, ensure the relevant data sets reflect characteristics specific to the geographic, contextual, behavioural or functional setting in which the high-risk AI system is intended to be used.
Official text
Article 10(4)Official source 4. Data sets shall take into account, to the extent required by the intended purpose, the characteristics or elements that are particular to the specific geographical, contextual, behavioural or functional setting within which the high-risk AI system is intended to be used. __________