In plain terms
Before a restaurant puts a new dish on the menu, someone checks the kitchen: are the ingredients in stock, are they fresh, are they labelled, and is the kitchen allowed to serve them? Data readiness is the same check for an AI project. It asks whether the information the system needs can actually be reached, may lawfully be used for this purpose, is up to date, is described well enough to be understood and is good enough for the job.
Why it matters
Data is where most of the time and money in an AI project goes, and where most delays begin. A readiness check in the first weeks turns an unpleasant surprise in month six into a planned cost. Readiness is always judged per use case: the same data can be ready for a monthly report and unfit for an assistant that answers customers. For generative AI the hard part is usually unstructured content and permissions. The limit: waiting until all company data is clean is the wrong lesson. Prepare what the chosen use cases need.
Example
A bank plans an assistant for 900 relationship managers and checks its five sources. Credit policies and product sheets pass all five tests. Customer files fail on access: a third are scans with no readable text. Call notes fail on permission: customers consented to service use, and this is a new purpose. Market reports fail twice: no owner and no dates. The bank launches in eight weeks with the two ready sources and funds the other three as a separate workstream.
Most often confused with
Data Readiness vs. Data quality
Data quality asks whether the data is correct: accurate, complete, consistent. Readiness asks whether it can be used for this purpose now. A flawless customer database that the AI system cannot reach, or that consent rules forbid it to use, is high in quality and unready. The reverse also happens: imperfect data is often good enough for a narrow task. Quality is one of five conditions.
Under the hood
A readiness assessment starts from the use case and works backwards: list the data it needs, then test each source against five conditions. Accessible: an interface or connector exists, formats are machine-readable, scans have been through OCR. Permitted: the legal basis and consent cover this purpose (GDPR, KVKK), licences allow it, access rights are correct and are carried through to retrieval, sensitive fields are masked. Current: the update frequency fits the task and expired content is archived. Documented: a named owner, a description of what each field or document means, lineage and validity dates. Quality: accuracy, completeness and duplicates, measured against what the task can tolerate. Predictive models add length of history, labels and representativeness. Generative AI adds parsing of tables and slides, metadata for filtering, and a test set of real questions with known answers. Typical findings: no owner, permissions wider than intended, and key knowledge held in people's heads or mailboxes. Outputs: a decision per source to use, fix or drop it, an effort estimate and named owners.