DatriseAI-first ETL

Upsert generator

Upsert into DuckDB: INSERT OR REPLACE / ON CONFLICT DO UPDATE

Name your table, key and columns and get the statement DuckDB actually accepts, with a guard so an older row never overwrites a newer one. Below it: how the mechanic works, what breaks, and how Datrise loads DuckDB incrementally.

Generate the statement

The URL updates as you type; share it to hand someone the exact form.

Statement · INSERT OR REPLACE / ON CONFLICT DO UPDATE

-- Target needs a PRIMARY KEY (or UNIQUE) on the conflict columns.
INSERT INTO "deals" ("id", "name", "stage", "amount", "owner_id", "updated_at")
SELECT "id", "name", "stage", "amount", "owner_id", "updated_at"
FROM "deals_staging"
ON CONFLICT ("id") DO UPDATE SET
  "name" = EXCLUDED."name",
  "stage" = EXCLUDED."stage",
  "amount" = EXCLUDED."amount",
  "owner_id" = EXCLUDED."owner_id",
  "updated_at" = EXCLUDED."updated_at"
WHERE "deals"."updated_at" IS NULL
   OR "deals"."updated_at" < EXCLUDED."updated_at";

-- Shorthand when every column should be replaced:
-- INSERT OR REPLACE INTO "deals" SELECT "id", "name", "stage", "amount", "owner_id", "updated_at" FROM "deals_staging";

How the upsert works in DuckDB

DuckDB implements the Postgres upsert syntax, INSERT … ON CONFLICT DO UPDATE with EXCLUDED, plus the shorter INSERT OR REPLACE INTO. Both need a PRIMARY KEY or UNIQUE constraint on the conflict columns, and DuckDB does not add one when you create a table with CREATE TABLE AS, so declare the table explicitly first. The staging source can be a file: read_parquet or read_csv_auto works directly in the SELECT.

A DuckDB database file has a single writer. That suits a local or notebook load where one process owns the file; it does not suit several workers upserting at once. For a file others read, generate a complete snapshot with CREATE OR REPLACE TABLE … AS SELECT, or export Hive-partitioned Parquet and let readers dedupe, which is the snapshot pattern Datrise uses for this destination.

Before you run it

  • DuckDB supports the Postgres ON CONFLICT syntax and the shorter INSERT OR REPLACE; both need a PRIMARY KEY or UNIQUE constraint, which DuckDB does not create by default on CREATE TABLE AS.
  • The staging "table" can be a file: FROM read_parquet('deals/*.parquet') or read_csv_auto(...) works directly in the SELECT.
  • A DuckDB file has one writer at a time. For a shared file, generate the full snapshot with CREATE OR REPLACE TABLE … AS and swap it, instead of concurrent upserts.

Questions people ask

INSERT OR REPLACE or ON CONFLICT DO UPDATE?

OR REPLACE overwrites every column of the matched row. ON CONFLICT DO UPDATE lets you list columns and add the watermark WHERE, so it is the safer default.

Why does DuckDB say the table has no primary key for the conflict target?

CREATE TABLE AS does not create constraints. CREATE TABLE deals (id VARCHAR PRIMARY KEY, …) first, then insert.

Can two processes upsert into the same .duckdb file?

No, one writer at a time. Run the loads sequentially, or write per-process Parquet and merge with a single DuckDB session.

The same generator for other destinations

Skip writing the merge at all

Datrise lands CRM and SaaS entities into DuckDB with this exact mechanic, a watermark on updated-at, and typed columns, so the statement above is what runs on your behalf. Join the waitlist to get early access.

Browse the integration catalog