All comparisons

IVYX vs Kiln AI and Transformer Lab

Kiln AI, Transformer Lab and IVYX Studio are all local desktop workspaces for AI work, and they overlap on evaluation, datasets and agents. The difference is what happens at the moment a model or an agent acts: Kiln AI and Transformer Lab record it, IVYX evaluates a rule first and can refuse. If your requirement is enforcement, and a signed evidence package proving it, that gate is the deciding feature. If it is not, Kiln AI and Transformer Lab are mature and very good at what they do, and on model training Transformer Lab is ahead of IVYX outright.

The short answer

  • Transformer Lab if the work itself is model training. It fine-tunes, including LoRA on diffusion models. IVYX does not fine-tune at all.
  • Kiln AI if you want a mature evaluation workflow with team ratings, synthetic data and Git-native dataset versioning, and you do not need enforcement.
  • IVYX Studio if someone outside your team will ask what ran, on what data, and who allowed it.
CriterionKiln AITransformer LabIVYX Studio
Local desktop appmacOS, Windows, LinuxmacOS, Windows via WSL2, LinuxmacOS and Linux
Fine-tuningYes, including distillationYes: full, LoRA and QLoRA, DPO, ORPO and SimPONo
Diffusion and imagesNoYes, including LoRA trainingNo
EvaluationLLM-as-judge, golden datasets, human ratingsLLM-as-judge, benchmarksDeterministic offline scoring on your golden sets
RAGYes, with chunking and indexingVia pluginsYes: chunking, embedding, indexing and retrieval
MCPYes, tools and MCP supportListed in the appYes, and it publishes servers as well as consuming them
Policy gate before executionNoNoYes
Human approval in the pathNoNoYes, for capabilities declared high-risk
Signed evidence packageNoNoYes, one per run
LicenceMIT library and server; source-available desktop appAGPL-3.0MIT source; the app runs on a 7-day trial, then a licence key

Recording and enforcing are different guarantees

All three keep a history of what ran. Kiln AI versions datasets through Git, Transformer Lab keeps runs in the app, and IVYX names every dataset and file its calls touched. Only IVYX evaluates a rule before the action and can refuse it, and only IVYX stops for a person on an action its manifest declares high-risk. A history assembled after the fact can prove harm occurred. A record built from gate decisions can show it was prevented.

Where IVYX is behind

On training, plainly. Transformer Lab does full fine-tuning, LoRA and QLoRA, preference-optimisation methods and diffusion LoRA. IVYX does none of that and has no plans to: fine-tuning orchestration is recorded as out of scope. Kiln AI's evaluation workflow, with team ratings and synthetic data generation, is more developed than IVYX's offline scoring. Both ship on Windows, which IVYX does not yet. And on price IVYX is the outlier of the three: Transformer Lab is AGPL-3.0 open source and the Kiln AI app is free to use, where IVYX Studio runs on a 7-day trial and then needs a licence key.

Frequently asked

Which local AI workspace has an audit trail?

Of the local desktop AI workspaces, IVYX Studio is the one built around an audit trail: every run records the datasets and files its calls touched, along with every gate decision and approval, as one signed evidence package. Kiln AI and Transformer Lab keep run history but do not produce an evidence package built from enforced decisions.

What is the best local alternative to Kiln AI?

Transformer Lab and IVYX Studio are the two closest local alternatives to Kiln AI. Transformer Lab goes deeper on model training, including diffusion and preference-optimisation methods. IVYX Studio adds a policy gate and a signed evidence package, which neither Kiln AI nor Transformer Lab provides.

Can I use Kiln AI or Transformer Lab alongside IVYX Studio?

Kiln AI, Transformer Lab and IVYX Studio can run side by side. All three are local desktop applications that work on ordinary files and datasets, so nothing prevents it. Training in Transformer Lab and promoting the result through IVYX Studio, so that the promotion carries a signed evidence package, is a reasonable split.

Competitor rows are drawn from each vendor's published material and were checked on 5 September 2026. A comparison page with a stale competitor row does more damage than no page at all, so if you are reading this much later, treat the competitor columns as dated.

Runs on your machine. macOS and Linux.