← JournalStudy No. 14

Development and testing

Checking reliability on new tasks

Diagnostics identified a systematic error. A partial correction helped on new individual tasks, including free generation. Complex sequences remain unresolved, and the criterion for proceeding to the next stage has not been met.

Status
Stopped
Outcome
Mixed
Bears on
a diagnostic observation
Period
4 — 7 October 2026
Published
Updated

What we set out to test

Does the improvement hold on new examples and when the model produces a complete answer on its own? Does it help on more complex tasks?

Earlier checks

How the study progressed

Events in order. The arrow shows which earlier event each one follows from.

  1. Preparation

    Starting a repeat study

    A repeat trial was prepared with initial-condition checks and retained results.

    Not assessed— bears on: a diagnostic observation

  2. Research

    Intermediate results

    ↳ Follows from: Starting a repeat study

    The first part of preparation passed its check. The next stopped at the session limit; there is no complete result yet.

    Not enough evidence yet— bears on: a diagnostic observation

  3. Diagnosis

    Identifying a systematic error

    ↳ Follows from: Intermediate results

    Related checks revealed a specific problem. Some other errors remained unexplained.

    Mixed— bears on: a diagnostic observation

  4. Fix

    Checking a partial correction

    ↳ Follows from: Identifying a systematic error

    The correction helped on some tasks. New examples were prepared to assess transfer.

    Mixed— bears on: a diagnostic observation

  5. Recheck

    The next-stage criterion remains unmet

    ↳ Follows from: Checking a partial correction

    Further training did not meet the required criterion. Progression to the next stage was stopped.

    Not enough evidence yet— bears on: a diagnostic observation

  6. Recheck

    The check on new tasks is complete

    ↳ Follows from: The next-stage criterion remains unmet

    Fully correct individual solutions in free generation increased from 10 to 19 out of 56. Complex sequences remain unresolved; no advantage over the simple comparison rule was established.

    Mixed— bears on: a diagnostic observation

What we found

The first part of preparation passed its check, but the next part has not reached the required criterion. Diagnostics identified a systematic error, followed by a test of a partial correction. The improvement held on new individual tasks. In a separate free-generation check, fully correct solutions increased from 10 to 19 out of 56. Complex sequences produced no fully correct solutions.

What the problem turned out to be

The partial improvement does not remove all errors. Some tasks still do not transfer to new examples. The improvement also did not establish an advantage for the choice under study over a simple comparison rule.

What we changed

We made a targeted correction and assessed it separately on new tasks.

What the recheck showed

The 7 October check is complete. A partial improvement was confirmed on this sample of individual tasks, including independently generated answers. Complex sequences remained unresolved. An advantage over the simple comparison rule was not established.

Limits of the conclusions

One experimental variant and small evaluation samples. The figures of 10 and 19 out of 56 apply to this particular check.

What is still unknown

  • Will the result hold in independent repetitions?
  • What prevents reliable solutions to complex sequences?
  • Will the approach prove more useful than simple comparison methods?

Next step

Investigate the remaining errors and test complete solutions to more complex tasks. The next stage begins once the set criterion is met.