Development and testing
Checking reliability on new tasks
Diagnostics identified a systematic error. A partial correction helped on new individual tasks, including free generation. Complex sequences remain unresolved, and the criterion for proceeding to the next stage has not been met.
- Status
- Stopped
- Outcome
- Mixed
- Bears on
- a diagnostic observation
- Period
- 4 — 7 October 2026
- Published
- Updated
What we set out to test
Does the improvement hold on new examples and when the model produces a complete answer on its own? Does it help on more complex tasks?
Earlier checks
- No. 12Preparing the current trialsNot enough evidence yet
- No. 13Checking multi-part tasksMixed
How the study progressed
Starting a repeat study
A repeat trial was prepared with initial-condition checks and retained results.
Not assessed— bears on: a diagnostic observation
Intermediate results
The first part of preparation passed its check. The next stopped at the session limit; there is no complete result yet.
Not enough evidence yet— bears on: a diagnostic observation
Identifying a systematic error
Related checks revealed a specific problem. Some other errors remained unexplained.
Mixed— bears on: a diagnostic observation
Checking a partial correction
The correction helped on some tasks. New examples were prepared to assess transfer.
Mixed— bears on: a diagnostic observation
The next-stage criterion remains unmet
Further training did not meet the required criterion. Progression to the next stage was stopped.
Not enough evidence yet— bears on: a diagnostic observation
The check on new tasks is complete
Fully correct individual solutions in free generation increased from 10 to 19 out of 56. Complex sequences remain unresolved; no advantage over the simple comparison rule was established.
Mixed— bears on: a diagnostic observation
What we found
The first part of preparation passed its check, but the next part has not reached the required criterion. Diagnostics identified a systematic error, followed by a test of a partial correction. The improvement held on new individual tasks. In a separate free-generation check, fully correct solutions increased from 10 to 19 out of 56. Complex sequences produced no fully correct solutions.
What the problem turned out to be
The partial improvement does not remove all errors. Some tasks still do not transfer to new examples. The improvement also did not establish an advantage for the choice under study over a simple comparison rule.
What we changed
We made a targeted correction and assessed it separately on new tasks.
What the recheck showed
The 7 October check is complete. A partial improvement was confirmed on this sample of individual tasks, including independently generated answers. Complex sequences remained unresolved. An advantage over the simple comparison rule was not established.
Limits of the conclusions
One experimental variant and small evaluation samples. The figures of 10 and 19 out of 56 apply to this particular check.
What is still unknown
- Will the result hold in independent repetitions?
- What prevents reliable solutions to complex sequences?
- Will the approach prove more useful than simple comparison methods?
Next step
Investigate the remaining errors and test complete solutions to more complex tasks. The next stage begins once the set criterion is met.