About this site 한국어

Unofficial explanatory translation of the Korean AI-Ready Data (AIRD) draft standard. The Korean text prevails.

Preparing data › Procedure

Checking your data

Type · ProcedureReading time · about 4 minData providers

ContentsContents
  1. Determining the data type
  2. What depends on the type
  3. Checking character encoding
  4. Data that contains personal information

Before starting stage 1, check the data type, character encoding and presence of personal information in the files you hold.

Determining the data type

The data provider decides the data type first. The measured items and metadata elements depend on the type. [Part 1 5.6 Table 5-5] Required

TypeDefinitionExamplesStandard identifier
TabularData with rows and columns, where each column has a defined data typeBusiness registration status · permit lists · transaction records · product master dataSTRUCT
Time seriesData that records values in time order. A special case of tabular dataEquipment sensor readings · daily pricesTSERIES
TextData made of natural-language sentences and documentsAnswers to civil complaints · original texts of public notices · reports · manufacturing technical documentsTEXT
ImageImage files such as photos and drawingsFacility inspection photos · satellite imagery · product defect inspection imagesIMAGE
Instruction–response pairsText data made of instruction–response pairs or preference pairsCustomer inquiry question–answer pairsPT

Data of multiple types Required

The data provider picks one representative type for the overall label. If a file or component falls under several types, apply the check items of each type to each component. Choosing a representative type does not mean skipping the check items of the other types. [Part 2 8.2]

Time series is a special case of tabular data. Apply the tabular check items to time series data as they are, and also record the reference time · measurement interval · unit. Recommended

What depends on the type

The data type determines the following three things.

ItemMeaningCount by typeBasis
Indicators that must be measured (required formal indicators)The set of indicators the standard specifies for each data type. Diagnostic Maturity (DM) reaches DM-2 only when every indicator in the set is measuredOf the 16 quality indicators: STRUCT 13 · TSERIES 13 · TEXT 5 · PT 8 · IMAGE 3[Part 2 Annex A.2]
Minimum measurable indicators (MMI)Indicators that must be measured at a minimum for any type. The lower-bound condition for the Diagnostic Maturity judgment4 — D5-03 (Statistical plausibility) · D6-01 (Uniqueness) · D7-01 (Encoding consistency) · D7-02 (Technical validity). Only the indicators that apply to the type are measured[Part 2 Annex A.1]
Additional metadata elementsElements added to the manifest depending on the typeSee Metadata elements[Part 3]

Required formal indicators define “how much must be measured to be complete.” Minimum measurable indicators define “what must be measured at the very least.” The rules that determine Diagnostic Maturity are in Judgment procedure; the definition of each indicator is in Quality indicators. Action items for text · image · instruction–response pair data will be included in a later version after they are checked against the text of the standard.

Without reference data to compare against, accuracy cannot be measured. The diagnostic tool marks this state as “accuracy not measured.” “Accuracy not measured” is a measurement status, not a penalty. If an accuracy indicator is a required formal indicator for that type, Diagnostic Maturity does not reach DM-2. In that case the move to the next readiness state (transition), that is, the transition to Quality-Ready, does not take place. [Part 2 7.4 · 8.1] Required

Checking character encoding

The reference encoding is UTF-8. CP949 · EUC-KR are not the reference encoding and are judged non-conforming. Encoding is measured by the minimum measurable indicator D7-01 (Encoding consistency). [Part 2 Annex A.4] The data provider can check and fix the encoding directly.

  1. Check with an editor. Open the file in a text editor and read the encoding indicator (status bar or the “Save As” dialog). The result is the name of the encoding.
  2. Check with a command. In a terminal, run file -i <file> (Linux) or file -I <file> (macOS). If the result contains charset=utf-8, the file is UTF-8. charset=us-ascii means the file contains only English letters and digits and can be read as UTF-8.
  3. Convert if it is not UTF-8. Save it again as UTF-8 in the editor, or run iconv -f CP949 -t UTF-8 <source> > <converted>. The result is a UTF-8 file.
  4. Record the result. Write down the encoding you checked together with the data type.

If your organization specifies a different encoding rule, the data provider follows it and records the applied rule in the manifest. Whether a byte order mark (BOM) is present and differences in line-break characters are not measured, so no action is needed.

Data that contains personal information

Check procedure when personal information is included

References: “Pseudonymized Information Processing Guidelines” (revised March 2026) of the Personal Information Protection Commission (PIPC, Korea’s data protection authority). Public institutions also check the “Public AX Privacy Guide” (July 2026). The latest versions are in the PIPC resource library.

A dataset that contains personal information must state whether it has been de-identified. [Part 3 Annex A.1] Required The level, method and date of processing are recorded in stage 2.

This procedure covers public provision. Data with restricted release is not covered by this procedure. Whether such data can be prepared in a controlled environment is judged separately by the organization.

Done when

  • You have decided and recorded the data type (one or more of STRUCT · TSERIES · TEXT · IMAGE · PT).
  • You have checked and recorded whether the character encoding is UTF-8.
  • You have checked whether the data contains personal information.

Writing the manifest (stage 1)

Last updated · 2026-10-07Report an error