Skip to content

SynapCores v2.0 — Run your transactions, your analytics and your ML on one engine

Stop paying for four systems to answer one question. v2.0 makes SynapCores a database you can put money through, point a BI tool at, and train models in — without moving a byte between vendors.


Put money through it: transactions that survive the power going out

The capability: financial-grade correctness for the workloads you currently keep in a separate OLTP database.

If your platform moves money, allocates inventory, or books capacity, "probably committed" is not a state you can reconcile in the morning. v2.0 gives you crash-durable transactions: when a COMMIT returns success, that data has reached the write-ahead log and will survive a hard power loss. Other connections cannot read your half-finished work, and SAVEPOINT lets a multi-step operation back out one stage without abandoning the whole thing.

What this retires: the reconciliation job that exists because you didn't fully trust the last write. And the second database you were running because this one couldn't be trusted with the ledger.

Proof: we kill the server with SIGKILL mid-transaction and restart it — zero partial rows. We acknowledge a 100-row commit, kill it, restart — all 100 rows present, with every constraint still enforced. 16 of 16 crash and concurrency gates pass on the shipped build.

Available on Community and Enterprise. No configuration required — it is the default.

[Read the durability guarantees →]


Stop bad data at the door: referential integrity that is actually enforced

The capability: your schema's own rules now do the work your application code was doing.

Declare a FOREIGN KEY and v2.0 enforces it — orphan rows are refused at insert, cascades run inside the same statement, and ON DELETE SET NULL does what it says. Table-level UNIQUE is enforced too, single-column and composite.

Why this matters commercially: every integrity rule that lives in application code is a rule each new service has to re-implement, and one of them eventually won't. Moving those rules into the database makes them true for every client, including the ones your team hasn't written yet.

Before you upgrade, run one query. Earlier versions accepted these declarations and never checked them, so you may already hold rows that violate constraints you believe are active. v2.0 refuses new violations but will not retroactively clean existing data — the audit query is in the upgrade guide. We would rather you find this on your terms than in production.

Available on Community and Enterprise.

[Get the pre-upgrade audit query →]


Answer analytics questions on your own storage, 2–3× faster than DuckDB

The capability: query Parquet in your own S3 bucket, in place, at speeds that make interactive exploration realistic.

What you're doing SynapCores v2.0 DuckDB 1.5.5
Scan 189,361 rows out of a lake table 10.1 ms 28.6 ms 2.8× faster
Count a join across two lake tables 5.5 ms 10.9 ms 2.0× faster
Join returning 176,400 rows 16.5 ms 29.1 ms 1.8× faster
GROUP BY one text column 5.0 ms 3.8 ms 1.3× slower

Same Parquet objects, same bucket, measured server-side on both engines, with every answer compared row by row. We publish the row we lose, because a benchmark that only shows wins tells you nothing about the one query you actually care about. If your workload is dominated by wide GROUP BY aggregation, DuckDB is still the faster reader for that shape.

What this unlocks: dashboards that aggregate lake data in milliseconds instead of seconds, and exploratory joins fast enough to actually explore with. No warehouse load step, no copy, no second vendor holding your data.

Available on Community and Enterprise.

[See the benchmark harness →]


Move production MySQL into your own S3, on a schedule, with proof it arrived

The capability: a managed export pipeline that replaces an ETL tool, a scheduler and the glue between them.

Point SynapCores at a production MySQL table, give it your bucket, and it writes Parquet to object storage on a cron you set — then registers that Parquet as an ordinary SQL table you can query in place.

The operational risk it removes: the 3am run that silently half-worked. Every run reports rows written, bytes, peak memory and the child process exit code, and verified_partitions tells you every file the job claimed to write was read back and confirmed. Credentials are encrypted at rest and the API never returns your secret — only the key id, so you can audit which credential is in use without being able to leak it.

Who this is for: teams running warehouse queries against a transactional database, where the fix has been "buy a warehouse" and the real problem is that nobody wants to own another pipeline.

Available on Community and Enterprise. Requires an admin to configure the destination credential.

[Follow the 20-minute walkthrough →]


Train, forecast and segment without moving data to a notebook

The capability: 32 machine-learning algorithms, trained with SQL, against the data already in your database.

v2.0 adds 19 algorithm paths, including two capabilities that were simply absent:

Find your customer segments without labelling anything. Clustering — KMeans, DBSCAN, Gaussian mixtures and more — needs no target column at all. Hand it behavioural columns and it finds the groups that are actually in your data, instead of the thresholds somebody guessed in a spreadsheet three years ago.

Forecast demand straight from SQL. Seasonal-naive, ETS, ARIMA, SARIMA and an additive changepoint model, with chronological validation — because a random train/test split on a time series quietly leaks the future and flatters the score.

Keep models current without owning an MLOps stack. A deployed model can retrain itself on a schedule that survives restarts, or when its inputs drift away from what it was trained on. Drift detection needs no labels — it compares live traffic against the distribution the model was fitted on — so it works on production data where the outcome isn't known yet. The industry answer to this is a monitoring library plus a separate orchestrator; here it is one clause in a SQL statement.

And money columns finally reach the model. A DECIMAL price or balance used to arrive as a missing value and get replaced by the column average — your model could not see spend at all. Fixed.

Honest scope: these are native implementations in the spirit of the well-known libraries, not those runtimes embedded, and they do not read or write upstream model files. Clustering quality is reported as a separation score, not an accuracy, because there is no ground truth to be accurate against. We would rather tell you that than let you present a silhouette as a 78% hit rate.

Available on Community and Enterprise. Optuna-based tuning requires a Python runtime on the host.

[Train your first model in SQL →]


Know when your database is actually unhealthy

The capability: a readiness probe that reflects whether SQL works, not just whether a process is alive.

A liveness check that only proves the HTTP listener is up will happily report green while queries hang — so nothing restarts, and nobody is paged. v2.0 adds a readiness endpoint that genuinely exercises the SQL path, and makes shutdown cooperative: on SIGTERM the engine stops accepting connections, drains the ones in flight, finishes accepted training work and exits cleanly, so a rolling deploy stops looking like a crash.

For platform teams: point livenessProbe at the cheap check and readinessProbe at the real one, and your orchestrator will finally take a wedged instance out of rotation.

Available on Community and Enterprise. Requires admin access to configure probes.

[Get the Kubernetes probe configuration →]


The trade-off we chose, and the number

Single-row writes are about 2× slower than v1.18. Roughly +0.7 ms on an individual INSERT, UPDATE or DELETE by primary key.

That is not a regression we missed — it is the cost of the durability guarantee at the top of this page. Every write now reaches the write-ahead log before it is acknowledged, which is precisely what makes an acknowledged commit survive a power cut. You cannot have the first property without paying for it.

Reads did not regress — every read case in our 39-case guard was within noise or faster, several lake shapes materially so.

If you are write-throughput-bound and don't need crash atomicity, one environment variable restores the previous write path. You give up isolation, crash atomicity and foreign-key enforcement to do it, which we think is the wrong trade for most teams — but it is your call, and it is documented.

Batching amortises it: wrap many rows in one transaction and the per-statement cost largely disappears.


Upgrading

Pull the new image or binary and restart. No schema migration, no client changes.

Three things to check first:

  1. Audit for constraint violations you already hold — FOREIGN KEY and table-level UNIQUE were not enforced before v2.0.
  2. Add retry handling for concurrent writers — real isolation means real conflicts. Catch the retryable error and retry the whole transaction. Code that never saw a conflict may start seeing them; that is the isolation working.
  3. If you send a bare BEGIN to the stateless query endpoint, it now fails loudly instead of silently autocommitting — which is the one change that can break working-looking code, though what that code was doing was never a transaction.

[Read the full upgrade guide →]


Known issues

  • Inside an explicit transaction, writes to a table carrying a trigger or a durable-agent binding are refused rather than silently excluded from the commit. Outside a transaction they work normally. A trigger's side effect is not part of the same commit as its row — full atomicity covers your rows and constraints, not hook effects.
  • AUTOML.PREDICT inside a common table expression requires the feature columns to appear in that CTE's own projection. The workaround is one line and is in the docs.
  • Two algorithm names are accepted by the parser for compatibility but have no implementation; asking for them is refused immediately with a list of what is available, rather than failing after you have waited.

Every number on this page was measured on the build this release ships, on a non-AVX-512 consumer CPU — not a tuned benchmark rig. Validation evidence, including the comparisons where we lose, ships in the repository.