About this site 한국어

Unofficial explanatory translation of the Korean AI-Ready Data (AIRD) draft standard. The Korean text prevails.

Preparing data › Procedure

Overview

Type · ProcedureReading time · about 4 minData providers

ContentsContents
  1. Stages and outputs
  2. What you can do now
  3. Five-question self-check of your current state
  4. Materials to gather before you start
  5. Before and after preparation
  6. Labels used on this site

This is a map of the procedure that checks the data you hold and, through three stages, brings it to a state AI can use. Data judged Quality-Ready or higher under an operating guideline is distributed as a pack; before that, you publish the manifest.

Scope — For tabular (STRUCT) and time series (TSERIES) data, this guide goes down to concrete action items. Text (TEXT) · image (IMAGE) · instruction–response pair (PT) data follow the same procedure. Items specific to each type will be added in later versions.

Stages and outputs

Data starts in a raw state, and its readiness state rises through three stages. Each stage ends at a decision gate. AI-ready data means a state of Quality-Ready or higher. [Part 1 3.1 · 5.1]

Three stages and decision gates
StageWhat you doOutputs (file · format)State reachedNeeded for the judgment
Checking your dataCheck data type · character encoding · presence of personal informationCheck recordRaw—
Stage 1 Writing the manifestWrite metadata elements · write the column schema · choose persistent identifiersManifest manifest.json — JSON-LD, discovery layer · column schemaDiscoverableRequired discovery-layer metadata elements met
Stage 2 Measuring quality · fixing defectsMeasure quality indicators · fix defects · record limitationsDiagnostic report — JSON · cleaned data · processing history · limitation record · understanding layer of the manifestQuality-ReadyThreshold profile of the operating guideline · Diagnostic Maturity (DM) DM-2 · quality tier Tier 1 or higher
Stage 3 Preparing for a purposeWrite operation files tailored to the use purposePurpose-specific operation file — one per purpose typePurpose-ReadyPurpose profile registered
DistributionAssemble and distribute the packPack — manifest · diagnostic report · pack manifest · operation files. Data files are referenced by address · checksum——

A new manifest is issued for each version and points to the previous version. The tier judgment procedure is in Judgment procedure; pack contents and file formats are in Packaging and distribution.

Working principles

  1. Follow the stage order. Make the data discoverable before you measure its quality.
  2. Complete what is not met. If data does not pass a decision gate, complete the unmet items and check again. Not meeting a gate does not mean rejection.
  3. Keep the original. Cleaned data and purpose-specific outputs are kept alongside the original.
  4. You can split the work by stage. Stage 1 alone reduces the items an agent must guess.

What you can do now

CategoryWork
A data provider can do aloneChecking the data · writing the manifest (23 metadata elements · column schema) · computing checksums. Choose persistent identifiers after checking your organization’s identifier policy
Needs a diagnostic toolMeasuring quality · writing the diagnostic report. The release version of the diagnostic tool is in development (Development status and upcoming releases)
Needs an operating guidelineQuality tier judgment · Quality-Ready transition judgment. The operating guideline provided with the standard is to be released when the standard is established (expected December 2026); organizations can judge with their own operating guideline

Details are in What is AI-ready data — What you can do now.

Five-question self-check of your current state

For data that is already open, the five-question self-check decides your starting stage. Depending on your answers, you start at writing the manifest (stage 1) or measuring quality (stage 2).

Judgment rule — the first “No” decides the starting point. An answer to a later question counts only after the conditions of earlier questions are met. Even if a later question is “Yes,” the starting point follows the first “No.”

The numbers in the “Choosing the starting point” table are the numbers of the five questions in the self-check above.

Choosing the starting point

First “No” Starting point Reason
Question 1 or 2 Writing the manifest (stage 1) Without an open format and column descriptions, quality cannot be measured
Question 3 or 4 Writing metadata elements in Writing the manifest (stage 1) These correspond to fields on the registration screen
Question 5 Measuring quality (stage 2) Recording limitations is stage 2 work
None (all “Yes”) Measuring quality (stage 2) Meeting the metadata elements and measuring quality are separate

The number of “No” answers does not indicate errors by the data provider. Existing portal registration screens had no input field for some of the 23 metadata elements and did not support delivery in machine-readable formats.

Materials to gather before you start

Before starting, the data provider gathers the following materials. If a material is missing, start with the metadata elements you can fill in.

Material Where to find it If missing
Data files The system in charge · portal registrations · internal data catalog Start with the description work in stage 1
Current registration information Public-sector data: portal registration screen · enterprise data: internal catalog or data exchange product listing Write it from scratch
Column descriptions Column comments in the source database Write them in Writing the manifest (stage 1)
Update frequency Actual update history Check the actual frequency first
Conditions of use Public-sector data: the institution’s open data policy · enterprise data: terms of use · license documents (with effective date) Ask the responsible department

Before and after preparation

Before and after preparation — expand 7-row table
Item Before After
Column meaning Column names only Name · data type · description · code values
Update frequency “As needed” URI from a standard list (e.g. …/frequency/IRREG)
Conditions of use “Includes third-party rights: N” License value from the vocabulary (e.g. KOGL-1)
File information None Access URL · media type · file format · checksum
Quality Unknown Scores for 7 dimensions + Diagnostic Maturity + evidence
Limitations None Bias · unsuitable uses · causes of missing values
History None What was changed · date · reason

Stage 1 does not change data values. Stage 1 adds information that describes the meaning, format and source of the values. Fixing defects in stage 2 can change values. The data provider records what was fixed, how and why in the processing history. Stage-by-stage results for an example file are in Walkthrough with example data.

Labels used on this site

LabelMeaning
RequiredRequired by the current draft standard
StandardSet by the operating guideline (threshold profile)
RecommendedRecommended by this site for convenience. Not a requirement of the draft standard
ExampleA value or wording for illustration

Recommended items are not requirements of the draft standard, so you can decide differently to suit your organization.

The mapping between these labels and shall · should · may in the specification text is not yet final.

Done when

  • You have answered the five-question self-check and decided your starting stage (stage 1 or stage 2).
  • You have checked whether you have each of the five materials in the “Materials to gather before you start” table.

Checking your data

Last updated · 2026-10-07Report an error