Insights Architecture
4 min read October 2026
What a modern data stack needs
Most teams need fewer tools than they think, and more of the unglamorous parts: tests, ownership, and documentation.
By the Flores DataCore insights desk

The phrase "modern data stack" usually arrives as a diagram covered in logos. The diagram rarely says what each box is for, what it costs to run, or who will be on call when it breaks. A better starting point is the work itself. Every data stack, whatever it is built from, does the same six jobs.
Six jobs, not six vendors
Get data in. Copy data from applications and databases into one place, either on a schedule or as it changes. Managed connectors are quick to set up; change data capture and custom code give more control.
Store it in one place. A cloud warehouse, a lakehouse, or, for modest volumes, one well-run PostgreSQL database. Keep raw data separate from cleaned and modeled data.1
Shape it. Apply business rules once, in version-controlled SQL: staging models that clean, core models that join, and marts that answer one team's questions.
Run it in order. An orchestrator or scheduler runs each step after the steps it depends on, retries what fails, and tells someone when it gives up.
Check it. Tests on keys, freshness, volume, and business rules catch problems before a stakeholder does.
Use it. Dashboards, forecasts, exports to other systems, and increasingly, machine learning.
Keep the layers apart
Raw data should land exactly as the source sent it, so history can always be replayed. Staging models rename and type columns without applying business rules. The rules live in the models above them. When a number looks wrong, that separation tells you where to look first: the source, the cleaning, or the rule.
It also makes change safer. A new source system replaces one staging model, not every report built on top of it.
What gets skipped
Stacks that disappoint rarely lack a tool. They lack the parts that do not come in a box.
- Ownership. Every important table has a named person who answers for it.
- Tests. A data test is a query that looks for rows breaking a rule. If it returns nothing, it passes.2
- Documentation. Plain-language definitions, kept next to the code they describe.
- Cost visibility. Knowing which jobs and dashboards drive the bill.
Right-size before you buy
If your data fits comfortably on one server, start there. A single managed PostgreSQL database, scheduled SQL, and one BI tool can carry a company a long way, and it is simple to understand. Distributed engines earn their keep at volumes and concurrency most small and mid-sized companies have not reached.
Concurrency matters as much as volume. Twenty analysts running heavy queries at nine in the morning stress a system differently than one nightly job does.
Pricing models matter as much as features. Many cloud services bill by use: by seconds of compute, by data scanned, or by rows synced.3 That is cheap to start and easy to lose track of. Before committing, estimate what happens to the bill if your data or your number of users doubles.
A minimal stack that holds up
- One ingestion path per source, with an alert when it fails.
- One storage platform, with raw, staging, and modeled layers kept apart.
- SQL models in version control, each with tests on its keys.
- One orchestrator or scheduler that knows the order of every step.
- One BI tool, reading only from modeled tables.
- A short glossary of the metrics leadership actually uses.
That is not glamorous. It is also the version that still works when the person who built it is on vacation.
Questions to ask before adding a tool
- Which of the six jobs does it do that nothing we already run can do?
- Who will own it, and what happens when they are out?
- How is it priced, and what happens to the bill if usage doubles?
- How would we leave? Open file formats and portable SQL keep that answer short.
- What will we stop using once it is in place?
The stack is the means. The point is a number someone trusts enough to act on.
- 1
Databricks calls this layering a medallion architecture: bronze for raw data, silver for validated data, and gold for enriched data. Other teams call the same layers raw, staging, and marts. Back
- 2
This is how dbt defines a data test: a query that selects failing records, which passes when it returns zero rows. dbt ships with four generic tests: unique, not_null, accepted_values, and relationships. Back
- 3
Examples from each vendor's documentation: Snowflake bills virtual warehouses in credits per second, with a 60-second minimum each time one starts; BigQuery's on-demand model bills by data processed, while its capacity model bills slot-hours; Fivetran prices connections by monthly active rows. Back

