About this site 한국어

Unofficial explanatory translation of the Korean AI-Ready Data (AIRD) draft standard. The Korean text prevails.

Preparing data › Procedure

Fixing defects (stage 2)

Type · ProcedureReading time · about 4 minData providers

ContentsContents
  1. Fixing low-scoring indicators
  2. Recording limitations that cannot be fixed
  3. Managing the original, the cleaned version and the diagnostic report
  4. Common errors

This page shows how to fix low-scoring indicators in stage 2 and how to record limitations that cannot be fixed. Keep the fixed file (the cleaned version) separate from the original.

GoalFix the values behind low-scoring indicators and record the limitations that cannot be fixed. Keep the original file, the cleaned version and the diagnostic report separate.

Fixing low-scoring indicators Recommended

The data provider’s cleaning work falls into three categories.

Category Work Dimensions improved
Cleaning column names Unify differing column names and notations into one Consistency (D2) · Machine readability (D7)
Cleaning values Align values with identifiers, classification schemes and standard codes Completeness (D1) · Consistency (D2) · Validity (D5) · Uniqueness (D6)
Adding descriptions Record measurement results and limitations in the manifest Machine readability (D7)

Priority actions by dimension

Priority actions by dimension — expand the 7-row table
Low-scoring dimension Priority action
Completeness (D1) Check empty values in key columns. If they cannot be filled, use one notation for empty values and record that fact
Consistency (D2) Check whether one concept goes by different names. Unify date and code notation
Accuracy (D3) Compare against the source material. If there is no reference material, leave it as “accuracy not measured” and record the reason
Timeliness (D4) Align the stated update frequency with the actual update history. State only a frequency you can keep
Validity (D5) Define rules for data types, allowed values and required values, and find rows that break them
Uniqueness (D6) Define the criterion for treating rows as the same (the record key). Without a key, record-level duplicate judgment is uncertain
Machine readability (D7) Convert to an open format. Remove merged cells, subtotal rows and multi-line headers

Recording limitations that cannot be fixed Required

The data provider also records limitations that cannot be fixed. Users read the recorded limitations to decide whether to use the data and how far. Users cannot see constraints that are not recorded.

What to record Example
Known bias Nationwide in scope, but over 90% of the data is from the Seoul metropolitan area
Unsuitable uses Statistics based on resident registration; unsuitable for analyzing the actual resident population
Causes of missing values Missing values caused by sensor maintenance (2.1%) and communication errors (1.1%)
Reasons for not measuring Accuracy not measured because there is no source material to compare against

Recording a limitation is not a judgment of a defect. It is information users need to judge the scope of use.

Managing the original, the cleaned version and the diagnostic report Required

The cleaned version keeps the dataset’s persistent identifier. A cleaned version is not a different dataset. The data provider keeps the original file, the cleaned version, the release version, the manifest, the diagnostic report and the purpose-specific operation files distinct from one another.

Distinguishing the original, the cleaned version and versions
What to distinguish How to distinguish it
Original file and cleaned version By distribution and SHA-256 checksum
Release version By version number
Manifest A new one is issued for each version and refers to the previous version [Part 3 5.1]
Diagnostic report A new one is produced for each measurement
Purpose-specific operation files A separate file is produced for each use purpose type

Keep the original file. Do not overwrite the original file with the cleaned version.

When an existing diagnostic report can be reused

Reuse an existing diagnostic report only when all four of the following are the same. If any one changes, produce a new diagnostic report.

  • The version measured
  • The measurement rules
  • The operating guideline applied
  • The validity of the measurement time

If the data was cleaned or measured again, produce a new diagnostic report. The manifest points to the new diagnostic report by its address and checksum.

Items to record in the processing history [Part 4 12]Follow-on part

Item Example
Work done Imputing missing values · normalizing code values · removing duplicates
Date of work 2026-08-14
Version change Original v1.0.0 → cleaned version v1.1.0
Reason for change Filled in missing district-level values · split age groups into 5-year bands

Without a processing history, users cannot verify the cleaning results.

Common errors

The three most common errors in stage 2 are listed below.

  • Judging the tier from the average score (tier judgment starts from the minimum dimension score)
  • Publishing a tier calculated with values that are not an operating guideline (for example, the example threshold profile)
  • Not updating the diagnostic report after cleaning
WrongCorrectedCause of the problem
Summarizing dimension scores D1 0.92 · D2 0.88 · D3 0.93 · D4 0.86 · D5 0.89 · D6 1.00 · D7 0.97 as a weighted average of 0.925 and judging the tier from itStart the candidate tier calculation from the minimum dimension score, D4 0.86An average dilutes a critical shortfall in one dimension. The weighted average 0.925 is calculated with the dimension weights of the example threshold profile (Part 2 Appendix I, status EXAMPLE). The example threshold profile is not used for tier judgment
Assigning a tier to the dimension scores of case J-GR-001 without an operating guidelineRecord only the scores and Diagnostic Maturity. The quality tier is “not judged: no operating guideline applied”The same scores yield different tiers depending on the operating guideline applied. A value calculated with the example threshold profile is not a tier
Leaving the previous diagnostic report in place after cleaning the dataProduce a new diagnostic report. The manifest points to it by its address and checksumAn existing diagnostic report is reused only when the version measured, the measurement rules, the operating guideline and the measurement time are all the same

Why it matters

The work in stage 2 creates the basis on which users judge whether the data can be reused. The data provider leaves limitations and a processing history along with the quality scores. Without the state of the values and the reasons for changes, users cannot judge the data’s constraints.

Quality scores and tiers are requirements the standard adds for AI use. The international FAIR principles do not require quality scores or tiers.

Completion criteria

  • Cleaned version, processing history and limitation records written
  • Original file kept (a different file with a different checksum from the cleaned version)

What you need Measure the cleaned version again with the diagnostic tool to produce a new diagnostic report. This requires the diagnostic tool. The release version of the diagnostic tool is in development (Development status and upcoming releases).

Reading the diagnostic report (stage 2)

Last updated · 2026-10-07Report an error