About this site 한국어

Unofficial explanatory translation of the Korean AI-Ready Data (AIRD) draft standard. The Korean text prevails.

Reference

Quality indicators

Type · ReferenceReading time · about 4 minAll readersAgent and AI developers (data users)

ContentsContents
  1. Application by data type
  2. Preconditions and formulas by indicator
  3. Calculation
  4. Partial measurement
  5. Exclusivity — no double counting of the same defect across indicators
  6. Reference encoding

This page lists how the 16 formal quality indicators apply to each data type, with their preconditions and formulas. All scores range from 0 to 1. Weights and tier bands follow the threshold profile of the operating guideline.

Application by data type

● required formal · ○ optional · — not applicable. MMI is a minimum measurable indicator (an entry requirement for diagnosis).

IDIndicatorDimensionSTRUCTTEXTIMAGETSERIESPT
D1-01Required field completenessD1 Completeness●○○●●
D1-02Required metadata completenessD1 Completeness●●●●●
D2-01Inter-field consistencyD2 Consistency●——●—
D2-02Referential integrityD2 Consistency●——●—
D2-03Derived value accuracyD2 Consistency●——●—
D3-01Standard name conformanceD3 Accuracy●○—○—
D3-02Label accuracyD3 Accuracy—○●—●
D4-01CurrencyD4 Timeliness●●○●●
D4-02Update fidelityD4 Timeliness○○——○
D4-03Time series completenessD4 Timeliness———●—
D5-01Format validityD5 Validity●●—●●
D5-02Numeric range validityD5 Validity●——●—
D5-03Statistical plausibilityD5 ValidityMMI●——●—
D6-01UniquenessD6 UniquenessMMI●○—●●
D7-01Encoding consistencyD7 Machine readabilityMMI●●○●●
D7-02Technical validityD7 Machine readabilityMMI●●●●●
Number of required formal indicators1353138

Preconditions and formulas by indicator

D1-01 Required field completeness — share of non-missing values in required columns

Precondition The list of required fields is defined in the column schema (csvw:tableSchema), in the metadata specification applied to the dataset, or in an equivalent document approved by the organization

Formula

(1 / |Freq|) × Σ_{f∈Freq} [ n(f ≠ NULL ∧ f ≠ '') / N ]

Symbols Freq set of required fields · |Freq| number of required fields · N total number of records

D1-02 Required metadata completeness — share of required metadata elements filled in

Precondition The metadata specification applied to the dataset names the list of elements that serves as the denominator of this indicator. For a specification that does not name such a list, all required elements of that specification form the denominator

Formula

n(required metadata elements filled in) / n(all required metadata elements)
D2-01 Inter-field consistency — compliance with relationship rules between two columns

Precondition The relationship rules are defined in one of the column schema (csvw:tableSchema), the metadata specification applied to the dataset, or an equivalent document approved by the organization

Formula

1 − ( n(records violating a rule) / N )
D2-02 Referential integrity — whether code values appear in the code table

Precondition Values are checked against the actual values of the master code table. If only an address or reference exists and the values cannot be obtained, the judgment is precondition unmet (PRECONDITION_UNMET)

Formula

1 − ( n(unmapped code cells) / n(valid code cells) )

Prerequisite indicator D5-01

D2-03 Derived value accuracy — whether stored derived values match the calculation rule

Precondition The calculation rule for derived values is defined in the schema, the column schema (csvw:tableSchema), or the metadata specification applied to the dataset. If there are no stored derived values, the judgment is not applicable (NOT_APPLICABLE)

Formula

1 − ( n(computed value ≠ stored value) / n(values checked) )
D3-01 Standard name conformance — whether codes and names match the authoritative reference

Precondition An authoritative reference is available as a code-to-name mapping, and the data contains both codes and names

Formula

1 − ( n(code cells not matching the reference) / n(matched code cells) )

Prerequisite indicator D2-02

D3-02 Label accuracy — share of labels matching the ground-truth labels

Precondition A set of ground-truth labels reviewed by experts exists. Automated measurement is not possible; the review results are used as input

Formula

n(matches) / n(items reviewed). For text and paired text, only an exact match counts as a match
D4-01 Currency — whether the data was updated within the stated update frequency

Precondition The update date and update frequency are recorded in the metadata

Formula

1.0 if days elapsed ≤ update interval. Otherwise max(0, 1 − (days elapsed − update interval) / grace period). The grace period equals the update interval in days
D4-02 Update fidelity — share of scheduled updates carried out

Precondition The update frequency and update history are stated in the metadata

Formula

n(updates carried out) / n(updates scheduled)
D4-03 Time series completeness — whether time series measurement points are missing

Precondition The expected measurement points are stated in the metadata

Formula

n(actual measurement points) / n(expected measurement points)
D5-01 Format validity — whether values follow the defined format

Precondition None. Columns without a defined format are excluded from measurement

Formula

1 − ( n(cells with format mismatch) / n(cells checked for format) )
D5-02 Numeric range validity — whether numbers stay within the allowed range

Precondition An allowed range is defined for each column. If no column has a defined range, the judgment is precondition unmet (PRECONDITION_UNMET)

Formula

1 − ( n(out-of-range cells) / n(cells with a defined range) )
D5-03 Statistical plausibility — whether dummy values or obviously wrong values are present

Precondition Dummy value patterns (indicator parameters in the threshold profile). Without an operating guideline, measurement is provisional using the example patterns in Part 2 Appendix I (Judgment procedure).

Formula

1 − ( n(cells with obviously wrong values) / n(numeric cells evaluated) )
D6-01 Uniqueness — whether records are duplicated

Precondition None

Formula

n(unique keys) / N. If identifier columns are defined, duplicates of the same entity are measured; if not, the combination of all columns is used as the key and fully identical duplicate records are measured
D7-01 Encoding consistency — whether file encodings match the reference encoding

Precondition None

Formula

n(files using the reference encoding) / n(all files)
D7-02 Technical validity — whether files open normally and have a valid structure

Precondition None

Formula

n(files that open normally ∧ meet the specification) / n(all files). A file meets the specification when parsing completes under the official specification of its file format and the required structure is valid

Calculation

ValueFormula
Dimension scoreDi = Σk (wk × Sk) / Σk wk
Minimum dimension scoremin(Di)
Weighted averageΣi (Wi × Di)
Minimum measurable indicator scoreΣ score(MMI) / n(MMI)
Weight redistributionWi′ = Wi / (1 − Wj) # Wj is the sum of the weights of the dimensions that could not be measured
Evaluation coverageΣ statusFactor(applicable required formal indicators) / n(required formal indicators excluding NOT_APPLICABLE)

Sk is the score of each required formal indicator for that data type whose measurement is complete. Optional indicators and incomplete indicators are excluded from both numerator and denominator. If a dimension has no completed formal indicator, Di is recorded as null. The weighted average is not used for tier judgment. Tier judgment starts from the minimum dimension score (Judgment procedure). Evaluation coverage is compared with the lower bound set in the operating guideline when judging Diagnostic Maturity DM-1.

Partial measurement

Indicators that allow partial measurement (PARTIALLY_APPLIED): D2-02 · D3-01 · D5-02. coverage = n(measurable targets) / n(identified measurement targets). Only targets that meet the precondition form the population. Targets that could not be measured are excluded from both numerator and denominator. The score is not multiplied by coverage again. If any measurement is partial, DM-2 cannot be reached. The upper limit of Diagnostic Maturity (DM) is DM-1.

Exclusivity — no double counting of the same defect across indicators

Defect typeResponsible indicator
Format violation in a single fieldD5-01
Out-of-range value in a single fieldD5-02
Relationship violation between two fieldsD2-01
Mismatch in a calculation across multiple fieldsD2-03
Empty cellD1-01
Unmapped codeD2-02

Reference encoding

ReferenceNon-conformingNot measured
UTF-8EUC-KR · CP949Presence of a byte order mark · differences in line break characters

[Part 2 Annex A.1–A.4] [Part 2 Annex B.1–B.6]

Last updated · 2026-10-07Report an error