Newport Resonance research
Multi-Step General Reasoning Without an LLM: A Structural-Mechanics Architecture Outperforms GPT-5.2 on a Seven-Family Reasoning Battery
A technical paper on structured multi-step reasoning, benchmark discipline, and what governed AI systems can learn from non-prompt-only evaluation methods.
- Status
- Published technical paper
- Publication date
- Author
- Ken Morkaya
Paper summary
Abstract
This paper examines whether multi-step reasoning can be evaluated through a structured control architecture rather than prompt-only language-model behaviour. The public abstract focuses on reasoning-task design, benchmark discipline, verifier-mediated repair, and the implications for governed AI systems where traceability and stability matter.
Result in context
What was measured—and what the comparison does not establish
Can support and contradiction be judged outside the language model?
Measure
Exact judgement rate on the paper’s direct-comparison sub-battery
Comparison · 40 author-generated evidence-judgement cases
- GPT-5.2 baseline
- 32/40 · 80%
- Structural reference
- 40/40 · 100%
- Scope
- This comparison covers the paper’s 40-case direct evidence-judgement sub-battery. It is not an overall comparison of the systems across every reasoning task.
- Study context
- The paper reports a 20 percentage-point observed difference on the 40-case direct-judgement comparison, with one GPT-5.2 run per case.
- Principal limitation
- The taxonomy and battery were author-designed, not independent. One model and one run per case were tested, and model-version or prompt-format variance was not measured.
Citation
Ken Morkaya. (2026). Multi-Step General Reasoning Without an LLM: A Structural-Mechanics Architecture Outperforms GPT-5.2 on a Seven-Family Reasoning Battery. Newport Resonance. /research/multi-step-general-reasoning-without-an-llm-a-structural-mechanics-architecture-outperforms-gpt-5-2-on-a-seven-family-reasoning-battery.pdf