Enunciado, How Faithfully Language Models Turn a Problem Statement into a Solvable Model
The field reports whether a generated model ran and calls that correct. Enunciado measures faithfulness instead: a corpus of 20 authored optimization statements across five complexity tiers, every trap covered and four controls without one, is handed to language models, and the formal model each returns is judged by an oracle that is not a language model: executable, structural and property layers over a solver, with a duality certificate. Sixteen models from four providers, 320 calls, every row complete. The best local model, phi4, is faithful on 5 of 20. The output cap decides the reasoning models: DeepSeek-V4-Pro is faithful on 2 of 20 at 8192 tokens and 11 of 20 at 32768. The first version of this measurement read +0.000 and was wrong.
Business Context
Anyone who hands an optimization, a simulation or an experiment design to a language model needs to know how often the returned model is the problem they stated, not merely a model that solves. Enunciado gives that number per model, per tier and per trap, with the caveats attached to the same ledger, so a team can pick a model for a formalization task on measured faithfulness rather than on a leaderboard that scored compilation. It also shows what moves the number: the output cap decides the reasoning models, the local lane silently shifts its default context, and a metamorphic refutation can be an unbounded candidate misread by three layers, each of them a recorded finding rather than a footnote.
Strategic Value
The product is the first adopter of a narrative-to-formal measurement archetype, and its record of what the gates caught is its strongest argument. Ten defects were caught by gates rather than by review: three wrong claimed optima; a solver wrapper that raised on the deliberately infeasible case instead of reporting infeasibility; two sweeps that shared one ledger and interleaved records from different code versions; a report that printed a gap of +0.000 when it meant no measurement; vendor names found in the CLI by the seam test; a deep link that answered 404 while rendering correctly. Four times a gate caught itself: the design-document guard read a prose cross-reference as a duplicate requirement, the UI gate passed its artifacts-loaded check while the page showed a router error, its theme switch wrote a key nothing reads so both passes ran in one theme, and its test server modelled the host wrongly in two different ways. Both Claude gaps of +0.050 rest on one refutation each that lands exactly on the reference whole-number optimum in a statement that never fixes integrality; read as allowed they are 0.000, the page states it, and the definition is a recorded decision. The other target families (mathematical and simulation modelling, experiment design, machine-learning framing), the manuscripts, and any measurement of a frontier model beyond the sixteen are not built and not claimed. Release 0.08.000 put the six pages in the standard order with App first, split the Benchmark into six tabs, and wrote the Spanish surface with its accents, 1,813 words decided in context.
The Challenge
Enunciado is the Spanish word for the statement of a problem: the text as it is handed to you, before anyone has decided what the variables are. Turning it into a formal, solvable model is the step the published evaluations skip. Four separate literatures report the artifact ran and call that correct, and each, read closely, admits the measured faithfulness is lower: in optimization modelling, reaching the reference objective does not imply a correct model (arXiv:2508.10047); in statement formalization there is a 3.0 to 29.0 point gap between compiling and being faithful, the strongest agent compiling 89.5 per cent of the time and being faithful 60.5 (arXiv:2606.31002); in experiment design every model is weak at datasets, baselines and metrics (arXiv:2608.03501); in simulation modelling the model runs but the causal reasoning is weak and no single model dominates (arXiv:2605.28994). A 2025 position paper argues these are one problem and supplies none of the machinery (arXiv:2509.09810). The community benchmarks cannot serve: they carry 8.13 to 54.0 per cent error rates, two of the most cited cannot be redistributed by licence, and the adjacent machine-learning family is contaminated.
Our Approach
Three repositories, one product. planteo (its own repository, on PyPI) is the representation: dimensions on every quantity, provenance on every element, and a record of what the statement left open. copela (its own repository, on PyPI) is the harness: a provider seam, an append-only ledger, four verdict layers and a budget guard that refuses to call a model without a declared budget. Enunciado is the product and declares no package: the corpus, the bake, the measurement and the web surface. The corpus is 20 authored cases, four per tier over five complexity tiers, every trap covered and four controls with no trap, written rather than imported for the reasons above. The bake verifies that every reference solves, every claimed optimum matches the solver and every property relation holds on the reference; it caught three wrong claimed optima out of twenty. The measurement calls sixteen models from four providers (Anthropic, Z.AI, DeepSeek and eleven open-weight models through Ollama on one 8 GB laptop GPU), 320 calls with every row complete in a 640-record ledger, and scores each answer through the executable, structural and property layers with a single failure taxonomy held equal to the classifier. The workbench shows fourteen methods in four groups, a Duality tab that is a four-part optimality certificate, a sidebar that diagnoses the selected case from the ledger, model matrices, a cap-sensitivity table, and caveats computed from the ledger rather than written.
Key Performance Indicators
| KPI | Baseline | Result | Impact |
|---|---|---|---|
| Faithfulness, not compilation | The published evaluations score whether the model ran | Executable, structural and property layers over a solver, with a four-part duality certificate; the best local model, phi4, is faithful on 5 of 20 cases | A model is chosen on what it got right, not on what it returned |
| The cap decides the reasoning models | A single number per model | DeepSeek-V4-Pro is faithful on 2 of 20 at an 8192-token output cap and on 11 of 20 at 32768; the cap-sensitivity table and the at-cap counts are published | A rank without its cap is not a result |
| A measurement that corrected itself | The first report read a gap of +0.000 | It meant no measurement; the report now distinguishes the two, and the 0.07.001 release re-derived report, attempts and rescore after CI found the artifacts covered 551 calls against a 640-record ledger | Every rate on the site is recounted from the committed ledger in CI |
Architecture
enunciado pipeline
Technology Stack
Application Screenshots

