About this site 한국어

Unofficial explanatory translation of the Korean AI-Ready Data (AIRD) draft standard. The Korean text prevails.

For agents

What agents read

Type · ExplanationReading time · about 5 minAgent and AI developers (data users)Tool developers

ContentsContents
  1. How the standard covers each judgment
  2. Proposal for exposing verified facts
  3. Summary of the coverage table

This page maps the information an agent needs for each of the four judgments it makes before using data to the clauses and metadata elements of the standard.

FitnessDoes it fit my question?Temporal coverageSpatial coverageUnitsPopulation coveredColumn meaningData typeUnsuitable usesReliabilityCan it be trusted?Actually updatedFile integrityAPI availabilityAPI responses match the specValues meet format · rangeDiagnostic MaturityJoinabilityCan it be joined with other data?Persistent identifierRecord keyStandard codes checked by valueEntity IDs and code list sourceUsabilityMay it be used?Machine-readable licenseAI training permittedAutomated collection permittedVersion of the terms of useFee · contract termsAccess conditionsContains personal dataRequiredPartial — cond. · opt. · rec., or not exposedGap — not in required elements or indicators
Coverage of the standard by judgment. Of the 24 pieces of information needed, 9 are required · 10 partial · 5 gaps. Joinability has no gaps. The 5 gaps are in fitness · reliability · usability.

Where the tiers agents read come from. The quality tier written in a manifest is judged not by the standard but by the threshold profile of an operating guideline. Whoever judged the tier records the operating guideline’s identifier · version along with the tier and publishes that operating guideline. When an agent reads a tier, it also reads which operating guideline was used to judge it. The conditions for using a tier as evidence are in Fields and handling by judgment.

The standard is not limited to public data. It covers all data published or provided by public institutions · enterprises · non-profit organizations. Its goal is to let AI agents access data, understand it, trust it, and judge whether they may use it. The design started from observations of public data. Agents judge from structured fields and measurement evidence, not from a single score. The “How the standard covers each judgment” table is based on tabular data (STRUCT). Documents · images · instruction · response pairs (PT) need different information for these judgments.

How the standard covers each judgment

Required required by the standard · Partial conditionally required · optional · recommended, or measured but not exposed as a field · Gap not in required elements or indicators (whether optional elements cover it is under review)

Of the 24 pieces of information needed: required 9 · partial 10 · gap 5

Fitness — fit to the question

Judgment question · Does it fit my question?

Fitness — fit to the question — expand 7-row table
Information neededStandardHow the standard covers it
Temporal coverageGapNone
The vocabulary term exists in DCAT-AP-KR (DCAT). It is not a required element of the AIRD discovery · understanding layers. Reference time and measurement interval for time series are recommended. Fill rate of “temporal coverage” in the portal catalog: 5.2% (Korea Public Data Portal catalog observation, 2026-08-31)
Spatial coverageGapNone
Same as temporal coverage. Fill rate of “spatial coverage” in the portal catalog: 6.4%
UnitsPartialUnit in the column schema (recommended) · time series unit (recommended)
Population coveredPartialPreparation item for the Stats (statistical analysis) purpose (recommended) · rai:dataBiases (known biases, required)
Column meaningRequiredColumn schema (csvw:tableSchema) (conditionally required for structured data)
Data typeRequireddct:type · aird:dataType
Unsuitable usesRequiredrai:dataLimitations

Reliability — grounds for trust

Judgment question · Can it be trusted?

Information neededStandardHow the standard covers it
Actually updatedPartialD4-01 Currency (required) · D4-02 Update fidelity (optional for STRUCT)
Update frequency is a required metadata element. Checking against the actual update history is an optional indicator
File integrityRequiredspdx:checksum · D7-02 Technical validity · suspension of validity (Stale)
API availabilityGapInterface conformance of the providing system is a level separate from data-side conditions [Part 1 Table 6-1]
The 16 quality indicators are file-based. Availability is not exposed as a fact attached to the data
API responses match the specGapNone
No API-side indicator corresponds to D7-02 for files
Values meet format · rangeRequiredD5-01 · D5-02 · D5-03
Diagnostic MaturityRequiredDiagnostic Maturity (DM) · 5 indicator applicability values

Joinability — ability to join

Judgment question · Can it be joined with other data?

Information neededStandardHow the standard covers it
Persistent identifierRequireddct:identifier
Record keyRequiredD6-01 Uniqueness (when identifier columns are designated)
Standard codes checked by valuePartialD2-02 Referential integrity · D3-01 Standard name conformance
The standard requires checking against the actual values of the code list (an address alone is not enough). Per-column results (code system and match rate) are only in the diagnostic report and are not exposed as metadata fields
Entity IDs and code list sourcePartialPreparation item for the KG (knowledge graph) purpose (recommended)

Usability — conditions of use

Judgment question · May it be used?

Usability — conditions of use — expand 7-row table
Information neededStandardHow the standard covers it
Machine-readable licensePartialdct:license (14 values in the vocabulary’s kr-license list — KOGL · CC · CC0)
The value list is closed. There is no value for an enterprise’s own terms of use or for MIT · Apache · ODbL licenses. Extending the value list is a candidate for a vocabulary revision proposal
AI training permittedPartialodrl:hasPolicy (optional) · KOGL AI type (kr-license)
The place for a machine-readable usage policy is an optional element. For data outside KOGL, no license value means training is permitted
Automated collection permittedPartialodrl:hasPolicy (optional)
Even data that people can download and use may have terms that prohibit automated collection (for example, web crawling). There are cases where the terms of data provided free by enterprises include such a clause
Version of the terms of useGapdct:rights notice (string)
Own terms change with each revision. No field records which version (effective date) of the terms applies
Fee · contract termsPartialdcatkr:fee (recommended) · sc:offers (conditionally required when fee=true)
Price · currency · unit can be recorded. There is no place for pay-as-you-go after a free quota or for terms limited to contracted customers
Access conditionsRequireddct:accessRights (eu-access: PUBLIC · RESTRICTED · NON_PUBLIC)
Contains personal dataPartialaird:deidentificationLevel (conditional) · dpv:hasPersonalData (conditional)

Proposal for exposing verified facts

Of the 7 facts in the “Standard status by fact” table, the standard already measures 5. The measurement results are recorded only in the diagnostic report. The manifest records only dimension scores. This site proposes a way to expose the measurement results in the manifest as facts.

Indicator measurementD2-02 Referential integrityTarget column sido_cdDiagnostic reportPer-column resultsCode list · version · match rateNowProposal — record as dqv:QualityMeasurementManifestDimension score Consistency 0.97ManifestIndicator D2-02 · target sido_cdBasis: admin. standard code (ver.)Value 0.997 · measured atValidity periodAgent“Consistency 0.97” alonecannot decide whether to joinAgentJoins on sido_cdCan cite the evidence
Same measurement, different exposure. D2-02 (Referential integrity) compares each column with its code list and computes a match rate. Currently that result is only in the diagnostic report, and the manifest records only the dimension score (top). The proposal records the result as one measurement (bottom). Example Values are illustrative.
Standard status by fact — expand 7-row table
FactJudgmentWhere the standard produces itStandard status
update_delay_daysReliabilityD4-01 (days elapsed − update frequency)Measured, not exposed
api_availability_7dReliabilityOutside the standardNot measured
response_matches_specReliabilityOutside the standardNot measured
pk_uniquenessJoinabilityD6-01Measured, not exposed
code_columns_matchedJoinabilityD2-02 (per column)Measured, not exposed
encodingReliabilityD7-01Measured, not exposed
required_fields_nonnullFitnessD1-01Measured, not exposed

Proposed exposure method

Proposal · not yet adoptedRecord verified facts in the manifest as dqv:QualityMeasurement. DQV is a vocabulary the standard already uses, so no new vocabulary is needed. One fact is one measurement and holds the indicator (dqv:isMeasurementOf) · value (dqv:value) · measured object (dqv:computedOn) · time (prov:generatedAtTime) · target column · reference basis (for example, a code list and its version) · validity period.

Summary of the coverage table

  • Joinability has no gaps. The standard requires checking against the actual values of the code list. What remains is exposing the per-column check results as metadata fields.
  • The 5 gaps are in fitness · reliability · usability. Temporal and spatial coverage are not required elements. Whether AI training is permitted appears only in an optional element. Agents use both kinds of information as conditions when selecting data.
  • Reliability indicators are file-based. The 16 quality indicators target files. Availability and spec conformance of data provided through an API do not appear as facts recorded with the data.
Last updated · 2026-10-07Report an error