Dependable data, clearly explained.

Data engineering & analytics

Flores DataCore, front page

Insights AI readiness

4 min read October 2026

Getting data ready for AI

Before a model comes history, labels, permissions, and a record of where every field came from.

By the Flores DataCore insights desk

Grid of wooden card catalog drawers with handwritten labels
Fig. 1A model can only learn from what has been filed, labeled, and kept. Note the drawer without a label.

When an AI project stalls, the cause is usually ordinary. The model is fine; the data is not ready. Nobody is sure there is enough history, nobody recorded the outcomes the model should learn, or nobody can say whether the data may be used this way at all. All of that can be checked before the first experiment.

Start from one use case

Readiness is not a property of your data in general. It is a property of the data for one decision. Predicting which invoices will be paid late needs invoice history across several billing cycles, the date each invoice was actually paid, and the facts that were known when the invoice was issued. A demand forecast needs something else entirely. Pick one use case and review the data for that.

Write down the decision the model will support, who will act on its output, and how you will know it worked. If nobody can answer those questions, the project is not ready, whatever the data looks like.

Four questions about the data

Is there enough history, and is it consistent? A new pricing plan, a changed definition of "active," or a system migration can split your history into periods that do not compare. Find those breaks before a model averages across them.

Do you know the outcomes? Supervised models learn from labels: the known result for past cases, such as paid late or paid on time. If the outcome was never recorded, or was recorded differently by different teams, that is the first job.

Is anything leaking from the future? A field like "days late" exists only after the outcome is known. Train on it and the model looks brilliant in testing and fails in use. Every feature needs an "as of" time, and the training set must respect it.

May you use it? Check what customers were told, what contracts allow, and which fields are sensitive. Get the data owner's agreement in writing, and record it next to the dataset. If a field is not needed, leave it out; data you never copy cannot leak.

Build features from the governed models

Feature tables should come from the same tested models as your reporting, not from one-off extracts. Then "active customer" in the model means exactly what it means on the dashboard, the same quality checks protect both, and lineage covers both. When a source changes, you know which models and which features feel it.

Hands holding a tablet that shows charts
Fig. 2The charts a forecast feeds should agree with the reports people already read.

Document the dataset

Researchers led by Timnit Gebru proposed that every dataset come with a datasheet describing its motivation, composition, collection process, and recommended uses.1 A one-page version covers most needs: what is in the dataset, where each field came from, the time range, known gaps, and who may use it for what.

Version and trace

Keep a snapshot or version of every training set, and record which model was trained on which version. With lineage from each field back to its source, a surprising result can be traced instead of guessed at, and a model can be retrained on the same data to check a fix.

Keep watching the inputs after launch. When the data feeding a model drifts away from the data it was trained on, accuracy usually falls before anyone notices.

The same discipline for generative AI

Retrieval-augmented generation looks up documents and passes them to a language model to ground its answer. The questions do not change: which documents, which versions, how fresh, and who may see them. Access rules have to carry through to retrieval, or a chatbot can surface a document its user was never allowed to open.

For governance, the NIST AI Risk Management Framework offers a voluntary structure, and its Generative AI Profile adds risks specific to generative systems.2

Readiness work is not glamorous. It is also the part of an AI project that carries over to the next one.

  1. 1

    Timnit Gebru and others, "Datasheets for Datasets," first posted in March 2018 and published in Communications of the ACM in December 2021. Back

  2. 2

    NIST released the AI Risk Management Framework (AI RMF 1.0) on January 26, 2023, for voluntary use, and the Generative Artificial Intelligence Profile (NIST AI 600-1) on July 26, 2024. Back