Skip to content
Research

Measured against a real toolchain.

An industrial engineering model is easy to make look good and hard to make correct. The difference only shows up against a compiler, a runtime and a plant's own project — so that is what SEIB is evaluated against.

No results published yet

SEIB has published no benchmark results, technical reports or model card.

The model is in training. When there are results, they will be published here, with the failure cases alongside them. Any figure attributed to SEIB before then did not come from SEIB — if you have been shown one, tell us.

Benchmarks

Four things worth measuring.

The public benchmarks for code generation do not cover industrial engineering, and the ones that come closest reward output that reads well. These are the axes SEIB is being built to be evaluated on.

01

Compile-and-run correctness

Does the generated artefact build in the toolchain it targets, and does it behave as specified in simulation? Pass or fail against a real compiler and runtime — not a similarity score against a reference answer.

02

Project-grounded accuracy

Given a real project's tags, routines and conventions, does the answer use that project's own vocabulary, or invent plausible names? This is where general-purpose models fail most expensively.

03

Cross-stage consistency

A change to the product changes the process, the automation and the commissioning. Measured end to end, because a model that is right at each stage and inconsistent between them has not helped anyone.

04

Refusal and escalation

How reliably the model declines work it should not do — anything touching safety-related program units above all — and hands it to an engineer instead of attempting it.

Technical reports

What we will publish, and when.

Written down now so it is a commitment rather than a plan. Nothing on this list exists yet; the first of it lands with the model, not before.

  1. 01A model card stating scope, intended use, and the work SEIB should not be pointed at.
  2. 02The evaluation methodology in full, including the harnesses and how the toolchain gates are run.
  3. 03Results against the benchmarks above, with the failure cases and not only the headline numbers.
  4. 04The safety evaluation: what is hard-blocked, how that is enforced, and how it was tested.
Experiments

The questions we are actually working on.

Open questions, described without results. If one of these is a problem you have data on or scars from, it is the most useful conversation we can have right now.

Verification in the training loop
Whether treating the compiler and the runtime as part of the training signal — rather than as a filter applied afterwards — measurably changes how often a first draft is usable.
Reading a plant as structure
Whether requirements, CAD and BOMs can be read as structured engineering objects rather than as documents to retrieve from, and what that changes downstream.
Carrying context across stages
How much of the intent behind a product change survives the trip from requirement to running line, and where in that chain today's tools lose it.

Working on the same problems? hello@seib.ai — or see who we are hiring.