AGENTIC DATA ENGINEERING

A coding agent specialized for data engineering

Vibedata runs the whole lifecycle above your data platform — with the isolation, guardrails, context, and cross-platform reach a general coding agent does not have. The model is a commodity dependency — the harness is the product.

Sandbox requests are open now; provisioning starts 10 August 2026.

Three agents with a chain of accountability: you direct the Build, Detect, and Fix agents inside the Vibedata boundary; work runs in code, compute, and data isolation, passes the gates, and one verified change lands on your data platform
THE PROBLEM

AI needs data faster than your team can supply it

Every agent and AI feature demands its own trusted data products, and the number a team must ship is exploding. But most data infrastructure wasn't built agentic-first — it was built for humans running dashboards and ad-hoc queries, not for agents that act on whatever they're given, at machine speed, with no error thrown when they're wrong.

Delivery hasn't kept pace either: a pipeline still takes two to four weeks to build, and teams lose about 40% of their time to reactive maintenance instead of shipping the next one.

Most teams do have guardrails for this. They are written as instructions: review before merging, never touch production, always test first. An agent that can read an instruction can also reason its way around it — and a data change that breaks nothing loudly gives no one a reason to look. Every published postmortem of an agent incident ends by shipping the same missing piece: somewhere the agent can be wrong on real data without production seeing it.

An agent working off broken data doesn't wait for someone to notice.

Supply versus demand: a small data team produces shared data products in the middle (one visibly incorrect) that many AI agents on the right consume
WHAT VIBEDATA IS

The world the agent works in

Vibedata is a coding agent specialized for data engineering, and the world around it — architecturally, a coordination layer above your data platform. That world is what a general coding agent lacks for data work: isolation, so being wrong is survivable; guardrails scoped to the data platform rather than to a generic external API; context that is data engineering rather than application code; and cross-platform reach. It is how agents build and maintain data products safely, on whatever platform you run. Every change is verified by independent execution against real data — the gates decide when work is done, and hallucination is structurally positioned out of the picture.

Our vision: every data team will be able to ship with the production discipline that used to take deep engineering expertise — specs, tests, isolation, and CI on every change — at vibe-coding speed.

Three agents run that lifecycle, with a chain of accountability:

  • The Build agent ships and modifies data products — ingestion and transformation together, from intent to a deployed pipeline.
  • The Detect agent watches the platform and reports problems and improvements, opening GitHub issues for the Build or Fix agent to act on — read-only, it never touches your code or data.
  • The Fix agent answers, investigates, and resolves — a fast, blast-radius-bounded loop with three depths: explain, investigate, remediate — and never merges its own fix: it asks before changing anything, and a human merges the PR.

It's built for two roles, not one: the AI-Native Data Engineer, who directs the Build agent on the build side, and the Data Reliability Engineer, who works with the Detect and Fix agents on the operate side.

What's yours, and what's ours

Everything in the first list you already own, or could assemble yourself — that is the point of it. The second list is what's left, and it is built per platform.

Yours, and standard

  • Your platform, your perimeter, your keys. Vibedata runs inside your own environment — Kubernetes or local Docker — and reads only metadata, code, and execution traces. It never reads your rows.
  • dlt for ingestion, dbt for transformation. Standard code you own and can walk away with, the same on every Adaptor — dbt models, dlt pipelines, and plain SQL that run anywhere. No proprietary formats, no lock-in.
  • Any model through the LiteLLM gateway, under your keys, with local and remote MCP into the systems you already run. Model choice is a routing decision, stage by stage: frontier models where judgment carries the load, cheaper tiers where the work is well-specified. The gates hold quality either way, so you ride the model price curve down.
  • Any platform, through a pluggable Adaptor. DuckDB (local and MotherDuck) and Microsoft Fabric today; Snowflake, then the Postgres Adaptor, Databricks, and BigQuery on the roadmap.

Ours, and native per platform

A branch of the data

A zero-copy clone of the lakehouse, wired into the agent loop — so a wrong attempt is disposable. That is what makes it safe to hand the agent real work rather than watching every change.

Gates and a deploy lock

Gates that verify every change by executing it on real data, and a deploy lock the agent can't talk its way past.

One durable goal

Intent — a goal that runs. One durable objective the agents plan, execute and verify across many steps, from build into operate.

Open at every level: any platform via Adaptor, any code (dbt, dlt, SQL), any cloud (your own tenant), and any model — the four plugging into Vibedata with no lock-in
OPEN BY DESIGN

Open at every level

The four things above are not a bundle you have to accept — each is a choice that stays yours: your platform through the Adaptor, your code in dbt and dlt, your cloud in your own tenant, and your model under your own keys.

The runtime sits inside your own environment and reads only metadata, code, and execution traces — never your rows. Row-level access stays with your agents, on your compute, using your approved model endpoints. Remove Vibedata and the dbt models, dlt pipelines, and SQL it produced keep running without it.

THE THESIS

Making the case in public, not just pitching it

Our thesis: agentic data engineering needs more than a chat window. It needs a coding agent that knows what data work is, and a world around it that makes the agent's mistakes survivable. We argue that in the open, three ways. We show what we're shipping and why it matters to AI-Native Data Engineers and Data Reliability Engineers. We argue about what it means to be an AI-agent-native data platform, backed by sources, not vibes. And we teach the engineering methods behind the harness, so you can apply the same thinking in your own stack whether or not you ever touch Vibedata.

Three ways we make the case in public: evidence of what shipped, a point of view, and a repeatable engineering method

Get notified on new Proof, Provocation, and Practice posts

Dogfood evidence, category argument, and transferable engineering method — straight from the team building the harness.