Reads the context
Does the model read the allowed sections before it answers, and does the answer reflect them?
Independent results are planned but not published yet. The benchmark will measure how models read Strap context, respect what it says, and propose updates worth keeping.
Three questions, scored per model, once the method is published.
Does the model read the allowed sections before it answers, and does the answer reflect them?
Do boundaries, preferences, and settled decisions hold across a multi-step task?
Are proposed changes narrow, durable, and accepted rather than rejected as noise?