About this site 한국어

Unofficial explanatory translation of the Korean AI-Ready Data (AIRD) draft standard. The Korean text prevails.

Preparing data › Examples

Public data examples

Type · ProcedureReading time · about 6 minData providers

ContentsContents
  1. Four examples
  2. Measured indicators before and after stage 1
  3. Defects revealed by measurement
  4. Catalog observation case
  5. Completion criteria

Measurements of four real public datasets show how stage 1 documentation increases the number of indicators that can be measured, and how results differ by data type.

Example Measurements were taken on 2026-10-06. Quality tiers were not judged because no operating guideline was applied (not judged: no operating guideline applied).

Four examples

Each example pairs one data type with one representative use purpose and is built from real public data. The types are STRUCT (tabular) · TSERIES (time series) · TEXT (documents) · PT (instruction · response pairs). The purposes are KG (knowledge graph) · Stats (statistical analysis) · RAG (retrieval-augmented generation) · FineTuning (fine-tuning).

ExampleType × purposeSourceSize
Commercial (business district) store information — Sejong sampleSTRUCT × KGSmall Enterprise and Market Service · Korea Public Data Portal (data.go.kr) 15083033 (202606 release)15,816 rows
Hourly national electricity demand (2025)TSERIES × StatsKorea Power Exchange · Korea Public Data Portal 150652668,760 time points
Article corpus of the Public Data Act and the AI Data Administration ActTEXT × RAGSaved PDF texts from the National Law Information Center of the Ministry of Government Legislation (current and forthcoming versions)4 versions · 209 articles · 673 units
Statutory interpretation question–answer pairs (51 cases) + comparison with 7 interpretations by the Korea National Data OfficePT × FineTuningMinistry of Government Legislation Open API target=expc (8,881 listed → 51 found by case title and full-text search) · target=kostatCgmExpc (all 7)51 cases → 45 pairs

Measured indicators before and after stage 1

Each example was measured three times.

  1. The original data was measured before stage 1.
  2. The same original data was measured after stage 1. In stage 1, the data provider recorded required fields · formats · value ranges · inter-field relations · derivation rules · code lists in the column schema (csvw:tableSchema).
  3. The cleaned version, with defects fixed, was measured.
ExampleRequired formal indicators① Original · before stage 1② Original · after stage 1③ Cleaned
STRUCT × KG135 / 1311 / 1311 / 13
TSERIES × Stats135 / 1310 / 1310 / 13
TEXT × RAG53 / 54 / 55 / 5
PT × FineTuning84 / 87 / 87 / 8

No data values changed between ① and ②. Yet the number of measured indicators increased, because the definitions recorded in stage 1 are the input to stage 2 measurement. Diagnostic Maturity (DM) by type is described in Reading the diagnostic report (stage 2).

Defects revealed by measurement

TypeDefectIndicatorHandling
STRUCTCommas in subcategory names were changed to semicolons (61 cells)D3-01 (Standard name conformance)Replaced with code list names → 1.0
STRUCT82 rows differ only in store number, with the other 38 columns identicalD6-01 (Uniqueness) 1.0 (by identifier)Handled as entity candidates in preparation for a purpose (KG)
STRUCTLegal-dong names omit the ri — one name, “Geumnam-myeon”, maps to 26 legal-dong codes—Legal-dong code used as the KG entity identifier
STRUCTValues of the Korean Standard Industrial Classification and legal-dong code lists could not be obtainedD2-02 (Referential integrity) partially measured (coverage 0.6 · score not recorded)Obtain the code list files and measure
TSERIESCP949 encodingD7-01 (Encoding consistency) 0.0Saved again as UTF-8 → 1.0
TSERIESTrailing empty row (366 rows on the portal · 365 actual rows)D1-01 (Required field completeness) below 1Empty row removed
TSERIESUnit appears only in the description text—Unit recorded in the column description
TEXTPDF — headers · footers mixed into the body, and the effective date appears only in one header lineD7-01 (judgment method under review)Split by article and recorded the effective date at the source location
TEXTCurrent and forthcoming versions of the same law—Versions distinguished with valid_from · valid_to
TEXTUpdate frequency “as needed”D4-01 (Currency) not measuredRecorded the update commitment as a frequency value
PT6 pairs of interpretations with identical content and different case numbersD6-01 1.0Only one of each pair kept in the operation file
PT7 empty values for the requesting agency nameD1-01 0.983 → 1.0Restored from the case title prefix, with the source recorded
PTA 2018 interpretation cites an article deleted in 2022—Marked as a “deleted article” and not linked to the current article

Defects can remain even when an indicator scores 1.0. D6-01 by identifier does not find rows that differ only in identifier but have the same content. The data provider handles such defects in preparation for a purpose (stage 3).

Catalog observation case

This is the result of observing the catalog metadata of one public dataset in the industry domain. The agency name is not shown. The subject is industrial complex company information (379 rows · 9 columns · CSV).

The basis for observation is the 19 required metadata elements of the deliberation draft (15 for the dataset + 4 for the distribution). The remaining 4 of the 23 in the vocabulary demo profile were added after the observation and were not classified.

CategoryCountElements
Recorded by the portal6title · publisher · file format · access URL · description · keyword
Needs format conversion3update frequency · media type · data service type
Must be newly written10dataset identifier · language · lifecycle status · creator · contact point · creation method · license · rights · access rights · checksum
Not classified (added after the observation)4dataset structure type · theme · landing page · maintaining department

If the portal provides column structure information, a draft column schema can be made without the file. In this case, drafts for the 9 columns were made from the observed column names · data types · counts of distinct values. The data provider then adds a description for each column.

The observation result also flags items that need a personal information check. In this case, the phone number column was flagged as an item whose disclosure scope must be confirmed with the responsible department.

Even for a dataset with catalog metadata on the portal, 10 elements must be newly written. Data that is already registered on the portal still needs separate preparation for AI use.

The catalog observation case covers only catalog metadata observation and the discovery layer check. Once the file is obtained, a diagnostic report is added.

Completion criteria

  • You found an example of the same type as your data.
  • In that example, you checked how many more indicators were measured after stage 1 documentation.
  • You checked whether your data has the same defects (for example, CP949 encoding · a trailing empty row · a unit only in the description text).

What you need: a diagnostic tool to measure indicators, and an operating guideline to judge quality tiers.

Next step: how to bundle and distribute the outputs is described in Packaging and distribution.

Last updated · 2026-10-07Report an error