A coding agent specialized for data engineering
Vibedata runs the whole lifecycle above your data platform — with the isolation, guardrails, context, and cross-platform reach a general coding agent does not have. The model is a commodity dependency — the harness is the product.
Sandbox requests are open now; provisioning starts 10 August 2026.
AI needs data faster than your team can supply it
Every agent and AI feature demands its own trusted data products, and the number a team must ship is exploding. But most data infrastructure wasn't built agentic-first — it was built for humans running dashboards and ad-hoc queries, not for agents that act on whatever they're given, at machine speed, with no error thrown when they're wrong.
Delivery hasn't kept pace either: a pipeline still takes two to four weeks to build, and teams lose about 40% of their time to reactive maintenance instead of shipping the next one.
Most teams do have guardrails for this. They are written as instructions: review before merging, never touch production, always test first. An agent that can read an instruction can also reason its way around it — and a data change that breaks nothing loudly gives no one a reason to look. Every published postmortem of an agent incident ends by shipping the same missing piece: somewhere the agent can be wrong on real data without production seeing it.
An agent working off broken data doesn't wait for someone to notice.
The world the agent works in
Vibedata is a coding agent specialized for data engineering, and the world around it — architecturally, a coordination layer above your data platform. That world is what a general coding agent lacks for data work: isolation, so being wrong is survivable; guardrails scoped to the data platform rather than to a generic external API; context that is data engineering rather than application code; and cross-platform reach. It is how agents build and maintain data products safely, on whatever platform you run. Every change is verified by independent execution against real data — the gates decide when work is done, and hallucination is structurally positioned out of the picture.
Our vision: every data team will be able to ship with the production discipline that used to take deep engineering expertise — specs, tests, isolation, and CI on every change — at vibe-coding speed.
Three agents run that lifecycle, with a chain of accountability:
- The Build agent ships and modifies data products — ingestion and transformation together, from intent to a deployed pipeline.
- The Detect agent watches the platform and reports problems and improvements, opening GitHub issues for the Build or Fix agent to act on — read-only, it never touches your code or data.
- The Fix agent answers, investigates, and resolves — a fast, blast-radius-bounded loop with three depths: explain, investigate, remediate — and never merges its own fix: it asks before changing anything, and a human merges the PR.
It's built for two roles, not one: the AI-Native Data Engineer, who directs the Build agent on the build side, and the Data Reliability Engineer, who works with the Detect and Fix agents on the operate side.
What's yours, and what's ours
Everything in the first list you already own, or could assemble yourself — that is the point of it. The second list is what's left, and it is built per platform.
Yours, and standard
- Your platform, your perimeter, your keys. Vibedata runs inside your own environment — Kubernetes or local Docker — and reads only metadata, code, and execution traces. It never reads your rows.
- dlt for ingestion, dbt for transformation. Standard code you own and can walk away with, the same on every Adaptor — dbt models, dlt pipelines, and plain SQL that run anywhere. No proprietary formats, no lock-in.
- Any model through the LiteLLM gateway, under your keys, with local and remote MCP into the systems you already run. Model choice is a routing decision, stage by stage: frontier models where judgment carries the load, cheaper tiers where the work is well-specified. The gates hold quality either way, so you ride the model price curve down.
- Any platform, through a pluggable Adaptor. DuckDB (local and MotherDuck) and Microsoft Fabric today; Snowflake, then the Postgres Adaptor, Databricks, and BigQuery on the roadmap.
Ours, and native per platform
A branch of the data
A zero-copy clone of the lakehouse, wired into the agent loop — so a wrong attempt is disposable. That is what makes it safe to hand the agent real work rather than watching every change.
Gates and a deploy lock
Gates that verify every change by executing it on real data, and a deploy lock the agent can't talk its way past.
One durable goal
Intent — a goal that runs. One durable objective the agents plan, execute and verify across many steps, from build into operate.
Open at every level
The four things above are not a bundle you have to accept — each is a choice that stays yours: your platform through the Adaptor, your code in dbt and dlt, your cloud in your own tenant, and your model under your own keys.
The runtime sits inside your own environment and reads only metadata, code, and execution traces — never your rows. Row-level access stays with your agents, on your compute, using your approved model endpoints. Remove Vibedata and the dbt models, dlt pipelines, and SQL it produced keep running without it.
Making the case in public, not just pitching it
Our thesis: agentic data engineering needs more than a chat window. It needs a coding agent that knows what data work is, and a world around it that makes the agent's mistakes survivable. We argue that in the open, three ways. We show what we're shipping and why it matters to AI-Native Data Engineers and Data Reliability Engineers. We argue about what it means to be an AI-agent-native data platform, backed by sources, not vibes. And we teach the engineering methods behind the harness, so you can apply the same thinking in your own stack whether or not you ever touch Vibedata.
Get notified on new Proof, Provocation, and Practice posts
Dogfood evidence, category argument, and transferable engineering method — straight from the team building the harness.