Roadmap

Benchmarks

Independent results are planned but not published yet. The benchmark will measure how models read Strap context, respect what it says, and propose updates worth keeping.

leaderboard · not yet publishedPreview

What the benchmark will measure

Three questions, scored per model, once the method is published.

1

Reads the context

Does the model read the allowed sections before it answers, and does the answer reflect them?

2

Respects what it says

Do boundaries, preferences, and settled decisions hold across a multi-step task?

3

Proposes updates worth keeping

Are proposed changes narrow, durable, and accepted rather than rejected as noise?