About this site 한국어

Unofficial explanatory translation of the Korean AI-Ready Data (AIRD) draft standard. The Korean text prevails.

About the standard

Why it is needed

Type · OverviewReading time · about 2 minAll readers

When people do not know something, they ask the data provider; agents instead guess the facts that are not recorded with the data. Observations of open data show this problem.

Observations of open data

We observed 84,271 file-data entries in the catalog of the Korea Public Data Portal (data.go.kr) (2026-08-31). In the catalog we found three values that block an agent’s judgment.

52.2%Share of file data whose update frequency is “as needed”. An agent cannot judge whether this data is current.
20.8%Share of file data whose next scheduled registration date has already passed. A machine has no evidence to check whether the schedule was kept.
0Number of fields, among the 36 fields of the catalog open-status listing, that hold a column’s unit · code scheme.

Source: observation of the Korea Public Data Portal catalog open-status listing (2026-08-31, 84,271 entries). Two DCAT files exported by the portal could not be read by an RDF parser.

For enterprise data, conditions of use are the first obstacle. We checked the terms of use of three Korean private-sector structured datasets; all three could be downloaded free of charge. However, all three restricted redistribution · automated collection through their own terms of use. None of the 14 license values in the vocabulary represents custom terms. The content of custom terms is written as text in rights (dct:rights), and how to record the license element is under review. People read the terms and judge, but an agent cannot confirm the conditions of use.

Cause of the problem

What is missing is not the data itself but the facts needed to use the data. Columns with the same name use different units, the same fact is reported several times, and the scope of the conditions of use is not written down. People fill in missing facts by asking the data provider. Agents cannot fill them in. See examples in Case videos.

What the standard does

  1. It has the facts agents use for judgment (meaning · unit · reference time · conditions of use) recorded in a machine-readable format in the manifest.
  2. It has quality measured with the same formulas, and the indicators measured and not measured recorded together.
  3. It specifies the tier system and judgment procedure. Threshold values are set by an operating guideline separate from the standard, and each tier must state the operating guideline applied.
  4. It has the grounds of each judgment recorded, so that an agent withholds rather than guesses when there is no evidence.

How to build an operating guideline is in Building an operating guideline.

Last updated · 2026-10-07Report an error