Plausible but unconservative
A classical buckling load used without a knock-down factor, or a linear result used beyond its range. Confident, neat and unsafe.
For AI labs & data teams
Original problems, verified solutions, reasoning traces, rubrics and agentic simulation tasks, written and checked by a computational mechanics PhD.
{
"type": "agentic · simulation in the loop",
"task": "Build and run a CalculiX model of a simply supported [0/90]s laminate under uniaxial compression. Report the critical buckling load N_x,cr.",
"tools": ["python", "calculix"],
"reference": "closed-form orthotropic plate solution",
"checks": ["mesh convergence < 1 %", "half-wave number", "units"],
"answer": { "rel_tol": 0.03 },
"grading": "automatic + expert rubric"
}
The problem
Fluent answers are easy to generate and hard to check. In engineering the dangerous errors are the plausible ones, so good training data targets them.
A classical buckling load used without a knock-down factor, or a linear result used beyond its range. Confident, neat and unsafe.
A swapped support condition or sign convention silently changes the answer by a factor. The reasoning looks fine; the number is not.
Mixed unit systems and unchecked orders of magnitude. Expert data teaches models to sanity-check results the way an engineer does.
Beam, shell or solid? Linear or nonlinear? Knowing when a model applies is the core of engineering judgement, and the hardest part to learn.
What I deliver
For supervised fine-tuning, reinforcement learning with verifiable rewards and held-out evaluation, in the schema your pipeline expects.
Original problems from undergraduate to research level. Final answers are checked by derivation and, where needed, by simulation.
Step-by-step solutions that state assumptions, sign conventions and units, and show the checks an expert makes along the way.
Grading rubrics with partial credit for open-ended answers, pairwise preference judgements and error taxonomies for model outputs.
Held-out, contamination-resistant problem sets with calibrated difficulty, including tasks current frontier models fail.
Agentic tasks in which a model must write, run and interpret finite element or numerical code, with automatic answer checking.
Line-by-line review of model reasoning, showing where and why it goes wrong, as input for targeted data.
Coverage
Mechanical, aerospace and civil engineering, from first-year statics to research-level stability and stochastic analysis.
Sample task
An illustrative graduate-level stability problem with its reference solution, grading rubric and the failure modes it is designed to catch.
A thin-walled aluminium cylinder with radius R = 500 mm, wall thickness t = 1 mm and length L = 1 m (E = 70 GPa, ν = 0.33) carries uniform axial compression between clamped ends. Estimate a design buckling load using the NASA SP-8007 knock-down approach.
Reference solution| Criterion | Pts |
|---|---|
| Classical buckling stress with correct dependence on ν | 2 |
| Correct cross-section and classical load | 1 |
| Knock-down factor applied, with correct φ | 3 |
| Design load within ±2% | 2 |
| Governing mode and validity checks stated | 2 |
Quality
Expert data is only valuable if it is right, unambiguous and new to the model. These rules apply to every task.
Problems are written from scratch rather than adapted from textbooks or the web, keeping them out of pre-training data.
Every final answer is reached by at least two routes (analytical, numerical or simulation) before delivery.
Answers come with units, tolerances and a rubric, so automatic and human graders agree.
Tasks are tested against current models and tuned so they discriminate rather than saturate.
Every task states its idealisations, so a correct answer is unambiguous and a wrong one is diagnosable.
Delivered in your schema (JSONL, Markdown with LaTeX, rubric sheets) or directly on your annotation platform.
Experience
Sander has worked on multiple AI training-data projects for several clients: writing original engineering-mechanics problems, reference solutions and rubrics, and reviewing model reasoning. This builds on ten years of research in computational mechanics.
Next step
Tell me about your pipeline, the domains you need and the volumes involved. A pilot batch is a good way to start.