Mesa-S1
A 1.2M-parameter model that builds CAD parts by operating FreeCAD's interface: toolbar buttons, dialog fields, dropdowns, OK and Cancel.
Mesa-S1 is an experiment in whether small, fast decision models can operate real application interfaces for computer-use agents: a planner decides what to build, and a tiny model does the clicking. It does everything Taiga-S1 does, but where Taiga-S1 picks FreeCAD commands and lets code fill in each dialog, Mesa-S1 works the dialogs itself. FreeCAD is the testbed; the next step is applications without a scripting API, read through the operating system's accessibility tree.
You give it a goal, an ordered list of features like "plate 40Γ30Γ12 β Γ4 hole at (8, 3) β linear pattern Γ3 over 20 mm β fillet the top edges". It builds the part step by step. At every step it reads FreeCAD's live state (the feature tree and selection, plus the interface: which toolbar commands are enabled, which task dialog is open, and the current value of every field in it) and picks the next element to act on:
| Element | What happens in FreeCAD |
|---|---|
cmd:PartDesign_Pad |
the toolbar button's action is triggered; FreeCAD opens its Pad dialog |
set:lengthEdit |
a value is typed into the dialog's length field |
opt:changeMode=Through all |
a dropdown entry is chosen |
toggle:checkBoxReversed |
a check box is clicked |
click:OK / click:Cancel |
the dialog is committed, or backed out of when it was opened by mistake |
- Tiny and fast. 1.2M parameters, trained from scratch, ~1 ms per decision on a CPU. No LLM, no vision model, no screenshots.
- Everything Taiga-S1 does, through the real interface. Same goals, same feature vocabulary, same test suites, at nearly the same accuracy: 99.75% of 800 held-out parts built correctly, 98.9% correctly and cleanly (no stray objects left), against 90β100% per suite for Taiga-S1.
- Handles longer goals than it trained on. Trained on goals of up to 5 features, it builds 99% of 11-feature goals (~80+ interface steps).
- Recovers from mistakes. With 20% of its actions replaced by random clicks, it notices the damage (cancels the wrong dialog, undoes the wrong feature) and still builds the right part 92β100% of the time. On long goals it often leaves something a random click created behind (55β74% clean, against 86β90% for Taiga-S1): the main gap to close.
Results
Built correctly and cleanly (no stray objects left in the document), with the "right part" rate in brackets where it differs:
| Goal | Mesa-S1 | With 20% random actions injected | Taiga-S1 (commands, for reference) | Taiga-S1, random actions |
|---|---|---|---|---|
| Parts like the training set (levels 1 / 2 / 3, up to 5 features) | 100 / 100 / 99% (100) | 99 / 98 / 85% (100 / 99 / 98) | 100% | 100 / 93 / 95% (100 / 98 / 98) |
| A held-out feature combination (never seen in training) | 99% (100) | 89% (100) | 90% | 95% (96) |
| Another held-out pairing | 100% | 95% (98) | 100% | 94% (95) |
| 6β7 features | 97% (100) | 74% (97) | 100% | 86% (88) |
| 8β9 features | 99% | 70% (95) | 100% | 87% (97) |
| 11 features | 97% (99) | 55% (92) | 100% | 90% (95) |
"Built correctly" means the model issued Done and the final solid matches the target exactly (volumetric IoU β₯ 0.99); "cleanly" means no stray objects remain. Each row is 100 fresh synthetic goals in FreeCAD 1.1, driven through the GUI. Per-step accuracy against the teacher's choices is 99.86% on held-out states.
Across the 800 clean episodes the model made no wrong decision that cost a part (both failures were FreeCAD crashes); it left a stray object behind in 7. FreeCAD crashes. Driven through its GUI, FreeCAD 1.1 segfaults in roughly 3β5% of long episodes, inside its own geometry and dialog code. The runtime recovers like FreeCAD's autosave: it restarts FreeCAD, replays the episode so far (episodes are deterministic) and retries the action. In the clean evaluation FreeCAD crashed 69 times, 55 were recovered, and the episodes it could not recover count as failures above. With random actions injected: 79 crashes, 67 recovered.
Usage
from freecad_s1.model.net import from_pretrained
from freecad_s1.rollout import Policy
model = from_pretrained("shhivv/mesa-s1")
policy = Policy(model, device="cpu")
probs = policy.score(state, goal, elements) # {element: probability}, best first
next_element = next(iter(probs))
What to pass in:
state: a snapshot of the FreeCAD session from the included UI runtime (freecad_s1.ui.session.UiSession, which runs inside the FreeCAD GUI), including the open dialog and its fields.goal: the ordered feature list, plus a rough size of the finished part (bounding box, volume). Estimates are fine.elements: the interface elements currently available, as returned byUiSession.valid_actions().
The returned probabilities are calibrated: a softmax temperature (T = 2.66, stored in config.json) was fitted on 16k on-policy states with 20% injected random actions, halving the calibration error (ECE 0.092 β 0.046). On those states, many of them mid-recovery, the model's top choice matches the teacher's 88.5% of the time; several recoveries are valid (e.g. Cancel vs. undoing a field change) but the teacher accepts only its own.
To watch it build a part in the FreeCAD GUI (code: github.com/shhivv/biome-s1):
FREECAD_S1_REPO=$PWD FREECAD_S1_UI=1 /Applications/FreeCAD.app/Contents/MacOS/FreeCAD scripts/freecad_gui_server.FCMacro &
python scripts/gui_demo.py --model release/mesa-s1 --level 3 --split iid --seed 7
How it was trained
- Data. 12k synthetic modeling sessions built through FreeCAD's interface in hidden FreeCAD instances, about 391k decisions. A scripted teacher labels the right next element at every step (any order where order doesn't matter, e.g. which dialog field first). Random mistakes are mixed into a quarter of the sessions so the model also learns to recover: cancel a dialog opened by mistake, undo a wrong feature.
- Training. Supervised training, then DAgger: the model drives FreeCAD on its own, the teacher labels every state it reaches, including after injected random clicks, and it retrains on its own mistakes.
- Architecture. Taiga-S1's architecture: a 3-layer transformer encodes the session state and the goal; each candidate element attends to the state and the active goal item and gets one score. An element is encoded by the command it stands for, the word pieces of its id, its role (command, field, dropdown entry, check box, button) and its live value (the number in a field, whether an entry is selected or a box checked).
Scope and limits.
- It covers FreeCAD PartDesign workflows: sketches (rectangle, circle, hexagon), pad, pocket, hole, revolve, linear and polar patterns, mirror, fillet, chamfer and shell.
- Mesa-S1 chooses which element to act on. The numbers typed into fields come from the goal (a parameter stage), as in Taiga-S1, and each numeric field tells the model whether it already holds that value.
- Picking faces and edges in the 3D view and drawing sketch geometry stay semantic actions (
canvas:β¦), since the widget tree cannot see the 3D view. - By default the interface is read from FreeCAD's Qt widget tree from inside the application, and actions are applied to those widgets. It also runs on the macOS accessibility tree, from outside the app: reading the interface that way, it built 15 of 15 test parts correctly; also clicking and typing through it (dropdown choices excepted), 14 of 15, the 15th lost to a FreeCAD crash (
scripts/mesa_ax.pyin the repo). The document state still comes from FreeCAD's API.
Citation
@misc{mesa_s1_2026,
title = {Mesa-S1: a small model that operates a CAD interface},
author = {Shanmugam, Shiv},
year = {2026}
}
- Downloads last month
- 39