For AI labs & data teams

Engineering mechanics data that holds up

Original problems, verified solutions, reasoning traces, rubrics and agentic simulation tasks, written and checked by a computational mechanics PhD.

task_agentic_031.json reference checked
{
  "type": "agentic · simulation in the loop",
  "task": "Build and run a CalculiX model of a simply supported [0/90]s laminate under uniaxial compression. Report the critical buckling load N_x,cr.",
  "tools": ["python", "calculix"],
  "reference": "closed-form orthotropic plate solution",
  "checks": ["mesh convergence < 1 %", "half-wave number", "units"],
  "answer": { "rel_tol": 0.03 },
  "grading": "automatic + expert rubric"
}
Illustrative agentic task record.

The problem

Why mechanics is hard for language models

Fluent answers are easy to generate and hard to check. In engineering the dangerous errors are the plausible ones, so good training data targets them.

Failure mode 01

Plausible but unconservative

A classical buckling load used without a knock-down factor, or a linear result used beyond its range. Confident, neat and unsafe.

Failure mode 02

Boundary conditions & conventions

A swapped support condition or sign convention silently changes the answer by a factor. The reasoning looks fine; the number is not.

Failure mode 03

Units & magnitudes

Mixed unit systems and unchecked orders of magnitude. Expert data teaches models to sanity-check results the way an engineer does.

Failure mode 04

Model-form judgement

Beam, shell or solid? Linear or nonlinear? Knowing when a model applies is the core of engineering judgement, and the hardest part to learn.

What I deliver

Data for training, evaluation and agents

For supervised fine-tuning, reinforcement learning with verifiable rewards and held-out evaluation, in the schema your pipeline expects.

01

Problems with verified answers

Original problems from undergraduate to research level. Final answers are checked by derivation and, where needed, by simulation.

02

Reference solutions & reasoning traces

Step-by-step solutions that state assumptions, sign conventions and units, and show the checks an expert makes along the way.

03

Rubrics & preference data

Grading rubrics with partial credit for open-ended answers, pairwise preference judgements and error taxonomies for model outputs.

04

Benchmarks & evaluation sets

Held-out, contamination-resistant problem sets with calibrated difficulty, including tasks current frontier models fail.

05

Simulation-in-the-loop tasks

Agentic tasks in which a model must write, run and interpret finite element or numerical code, with automatic answer checking.

06

Expert review of model output

Line-by-line review of model reasoning, showing where and why it goes wrong, as input for targeted data.

Coverage

Solid and structural mechanics, end to end

Mechanical, aerospace and civil engineering, from first-year statics to research-level stability and stochastic analysis.

  • Statics & equilibrium
  • Strength of materials
  • Beams, torsion & stress transformation
  • Elasticity & continuum mechanics
  • Plates & shells
  • Buckling & stability
  • Post-buckling & imperfection sensitivity
  • Composite laminates
  • Failure criteria
  • Fatigue & fracture
  • Structural dynamics & vibration
  • Finite element method
  • Numerical methods & solvers
  • Uncertainty quantification & reliability
  • Surrogate modelling & ML for mechanics
  • Aerospace structures & loads
  • Additive manufacturing

Sample task

What a verified task looks like

An illustrative graduate-level stability problem with its reference solution, grading rubric and the failure modes it is designed to catch.

Problem

A thin-walled aluminium cylinder with radius R = 500 mm, wall thickness t = 1 mm and length L = 1 m (E = 70 GPa, ν = 0.33) carries uniform axial compression between clamped ends. Estimate a design buckling load using the NASA SP-8007 knock-down approach.

Reference solution
  1. Classical buckling stress: σcl = Et / [R√(3(1 − ν2))] ≈ 85.6 MPa
  2. Classical load: Pcl = 2πRt·σcl ≈ 269 kN
  3. Knock-down factor: φ = √(R/t)/16 ≈ 1.398 and γ = 1 − 0.901(1 − e−φ) ≈ 0.322
  4. Design load: Pd = γ·Pcl ≈ 86.5 kN
  5. Checks: the Euler column load (≈ 2.7 × 105 kN) is far above Pcl, so shell buckling governs; R/t = 500 lies within the method’s range; Batdorf parameter Z ≈ 1.9 × 103 (long-cylinder regime).
Pd ≈ 86.5 kN tolerance ±2% · unit kN
Rubric · 10 points
CriterionPts
Classical buckling stress with correct dependence on ν2
Correct cross-section and classical load1
Knock-down factor applied, with correct φ3
Design load within ±2%2
Governing mode and validity checks stated2
Failure modes it catches
  • Reports the classical load (≈ 269 kN) as the capacity: unconservative by a factor of about 3.1.
  • Applies the Euler column formula to a thin shell.
  • Uses the knock-down formula outside its range of validity, or drops units.

Quality

Quality, by construction

Expert data is only valuable if it is right, unambiguous and new to the model. These rules apply to every task.

  1. Original authorship

    Problems are written from scratch rather than adapted from textbooks or the web, keeping them out of pre-training data.

  2. Independent verification

    Every final answer is reached by at least two routes (analytical, numerical or simulation) before delivery.

  3. Unambiguous grading

    Answers come with units, tolerances and a rubric, so automatic and human graders agree.

  4. Calibrated difficulty

    Tasks are tested against current models and tuned so they discriminate rather than saturate.

  5. Stated assumptions

    Every task states its idealisations, so a correct answer is unambiguous and a wrong one is diagnosable.

  6. Your format

    Delivered in your schema (JSONL, Markdown with LaTeX, rubric sheets) or directly on your annotation platform.

Experience

Hands-on with AI training data

Sander has worked on multiple AI training-data projects for several clients: writing original engineering-mechanics problems, reference solutions and rubrics, and reviewing model reasoning. This builds on ten years of research in computational mechanics.

Scale
From pilot sets to ongoing batches
Pricing
Per task, per batch or hourly
Confidentiality
NDAs and data-handling requirements respected; tasks are never published

Next step

Building a mechanics dataset or benchmark?

Tell me about your pipeline, the domains you need and the volumes involved. A pilot batch is a good way to start.