About this site 한국어

Unofficial explanatory translation of the Korean AI-Ready Data (AIRD) draft standard. The Korean text prevails.

For agents

Fields and handling by judgment

Type · ProcedureReading time · about 4 minAgent and AI developers (data users)Tool developers

ContentsContents
  1. Fitness — fit to the question
  2. Reliability — grounds for trust
  3. Joinability — ability to join
  4. Usability — conditions of use
  5. Handling information the standard does not define
  6. Automated-use eligibility conditions

This page shows in pseudocode which fields an agent reads from the manifest for each judgment and how it judges them. The fields and judgment values follow the standard; the judgment rules are this site’s recommendations.

Reading principle — do not replace a missing field with a guessed value. If there is no evidence, the agent states “no evidence” in its result. When a human judgment is needed, the agent hands the case over to human review. The standard aims to let agents judge without guessing, and the judgment rules on this page follow the same principle.

Fitness — fit to the question

Judgment question · Does it fit my question?

Fields read dct:description · csvw:tableSchema · dct:type · rai:dataLimitations · rai:dataBiases

schema = m.get("csvw:tableSchema")
if schema is None:
    return no_evidence("column meanings are not recorded")   # do not guess from column names
for col in columns_used_by_question:
    if not column(schema, col).get("dc:description"):
        return no_evidence(f"meaning of {col} is not recorded")
# rai:dataLimitations · rai:dataBiases are lists of human-written sentences.
# Do not filter them by string matching; present them as-is with the answer.
caveats = m.get("rai:dataLimitations", []) + m.get("rai:dataBiases", [])
return fit(caveats=caveats)

When the field is missing If the column schema (csvw:tableSchema) has no column description, the agent does not guess the meaning from the column name. It states “column meaning not recorded” in the answer.

Reliability — grounds for trust

Judgment question · Can it be trusted?

Fields read aird:lifecycleStatus · dcat:distribution[].spdx:checksum · dct:modified · dct:accrualPeriodicity · aird:diagnosticMaturity · aird:qualityTier · aird:thresholdProfile · aird:diagnosticReport

if m["aird:lifecycleStatus"] == "Stale":
    return exclude("validity suspended")
d = distribution(m, downloaded_url)               # each distribution has its own checksum
if sha256(downloaded_file) != d["spdx:checksum"]["spdx:checksumValue"]:
    return exclude("checksum mismatch — not the file the manifest points to")
dm = m.get("aird:diagnosticMaturity")
if dm is None:
    # null has two meanings. Tell them apart by the diagnosis status in the diagnostic report.
    if diagnosis_status(m) in ("SCHEMA_ONLY", "FILE_BROKEN"):
        return exclude("file could not enter diagnosis")
    return caution("Diagnostic Maturity not established — insufficient quality evidence")   # do not read as 0
# A tier is judged by the threshold profile of an operating guideline. Check all four conditions.
profile = threshold_profile(m.get("aird:thresholdProfile"))   # None if absent
guideline = operating_guideline(m)                 # identifier · version · URL — None if absent
tier_evidence = (
    dm == "DM-2"                                              # (a) Diagnostic Maturity DM-2
    and profile is not None
    and profile["status"] == "OFFICIAL"                        # (b) EXAMPLE is not tier evidence
    and evaluation_status(m) != "PROVISIONAL"                 # (c) not a provisional evaluation
    and guideline is not None and guideline.identifier and guideline.version
    and accessible(guideline.url)                             # (d) the published operating guideline can be checked
)
if not tier_evidence:
    return caution(f"Diagnostic Maturity {dm} · no tier evidence", score=m.get("aird:qualityIndexMin"))
return trusted(tier=m.get("aird:qualityTier"),
               guideline=(guideline.identifier, guideline.version),
               profile=(profile["profileId"], profile["version"]))

When the field is missing If Diagnostic Maturity (DM) is missing, the agent treats it as “insufficient measurement evidence”, not as “low”. The agent reads a tier as evidence only when all four of the following conditions are met.

ConditionValue checked
(a) Diagnostic MaturityDM-2
(b) Status of the threshold profileOFFICIAL (not EXAMPLE)
(c) Evaluation statusNot PROVISIONAL
(d) Operating guidelineThe identifier · version of the operating guideline used for the judgment are known, and the published document is accessible. The standard does not yet have a manifest property that points to the operating guideline, so this is checked through the threshold profile identification (aird:thresholdProfile) and the operating guideline information recorded alongside the tier

If any one condition is not met, the agent sets the result to “caution (… no tier evidence)” and reads the score only for reference. The structure and publication requirements of operating guidelines are in Building an operating guideline.

Joinability — ability to join

Judgment question · Can it be joined with other data?

Fields read dct:identifier · csvw:tableSchema.primaryKey · csvw:tableSchema.columns[].dc:description · D2-02 result in the diagnostic report

a, b = join_columns_of_the_two_datasets
# The manifest has no dedicated field for a column’s code system (today it appears only in the column description text).
# Value-level check results are in the diagnostic report and are not exposed in the manifest — the verified-facts proposal fills this gap.
ra, rb = d2_02_result(report(a)), d2_02_result(report(b))   # None if the report cannot be read
if ra is None or rb is None:
    return candidate("value-level check results unavailable — do not join just because names match")
if ra.applicability != "APPLIED" or rb.applicability != "APPLIED":
    return candidate("values were not checked against a code list")
if ra.code_list != rb.code_list:
    return exclude("checked against different code lists")
return join(evidence=[ra.match_rate, rb.match_rate, ra.code_list_version])

When the field is missing If there is no value-level check result from D2-02 (Referential integrity), the agent keeps the join only as a “candidate”. It states in the result that there is no evidence for the join.

Usability — conditions of use

Judgment question · May it be used?

Fields read dct:license · dct:rights · dct:accessRights · dpv:hasPersonalData · aird:deidentificationLevel · odrl:hasPolicy

if m["dct:accessRights"] != "PUBLIC":
    return human_review("access-restricted data")
lic = m["dct:license"]
if lic not in KR_LICENSE:                      # value list kr-license in the vocabulary
    return human_review("license outside the value list — machines cannot read its conditions")
conditions = license_conditions(lic)
# No required element covers AI training permission — read only the optional odrl:hasPolicy; do not guess from dct:rights text
policy = m.get("odrl:hasPolicy")
training_allowed = permitted_actions(policy) if policy else "unconfirmed"
return usable(conditions=conditions, training_allowed=training_allowed)

When the field is missing Whether AI training is permitted is not a required element of the standard. It can be expressed in the optional usage policy (odrl:hasPolicy). If there is no usage policy, the agent does not guess from the rights (dct:rights) text and leaves it as “unconfirmed”.

Handling information the standard does not define

The following four kinds of information are not required elements of the standard. This site recommends the following handling on the agent side.

Information not in the required elementsRecommendation for agents
Temporal coverage · spatial coverageRead DCAT dct:temporal · dct:spatial if present. Otherwise, mark “coverage not recorded”. If the question requires coverage, hand over to human review.
Whether AI training is permittedIf there is an ODRL policy (odrl:hasPolicy), read the permitted · prohibited actions. Otherwise, mark “unconfirmed”. For training use, hand over to human review.
Whether automated collection is permitted · version of the terms of useIf the ODRL policy has a prohibited action (odrl:prohibition), follow it. If there is no policy and the conditions of use exist only as a terms document, do not collect automatically and mark “unconfirmed”. If the effective date of the terms cannot be confirmed, hand over to human review.
API availability · response conformance to the specThe agent checks these directly before calling. It keeps the check time and result as evidence for its judgment. It does not substitute values from the manifest.

Automated-use eligibility conditions

Part 5follow-on part defines 8 data-side conditions for automated-use eligibility. The pseudocode below covers all 8 conditions. Conditions for which the standard has not yet defined fields (2 · 3 · 4) are marked “under review”. Interface · exchange conformance of the providing system is handled at a separate level.

# Automated-use eligibility — 8 data-side conditions [Part 5 Table 9-2]follow-on part
# Only conditions whose fields the standard defines are written as code. Conditions with unconfirmed fields are marked.
def eligible_for_automated_use(m, file, url):
    d = distribution(m, url)
    return all([
        m["aird:lifecycleStatus"] == "Purpose-Ready",           # 1 Purpose-Ready data
        currently_provided(m),                                   # 2 field under review
        officially_published(m),                                 # 3 field under review
        purpose_tier(m) is not None,                             # 4 property name under review — aird:purposeTier / aird:purposeReadiness
        m.get("aird:diagnosticMaturity") == "DM-2",             # 5 all required items measured
        quality_tier_at_least_minimum(m),                        # 6 at least the minimum tier set by the operating guideline
        sha256(file) == d["spdx:checksum"]["spdx:checksumValue"],  # 7 checksum matches (per distribution)
        m.get("dct:license") in KR_LICENSE and m.get("dct:rights") and m.get("dct:accessRights"),
                                                                # 8 conditions of use can be confirmed — because this is
                                                                #   automated use, read as a machine-readable license
    ])

The “under review” marks in the pseudocode are where property names · interpretations are being checked against the original text of the standard.

The shape of a manifest is in Manifest example, and the list of properties is in Metadata elements. The release status of schema files and validation tools is in Development status and upcoming releases.

Other paths · Preparing data — Writing the manifest (stage 1) · Operating guidelines — Adoption brief

Last updated · 2026-10-07Report an error