QuickScope

Average scores hide the interesting failures.

On saturated benchmarks, strong models can look nearly tied on headline averages. QuickScope is aimed at the diagnostic question underneath: where does a model still fail reliably, and what kind of benchmark region produces that failure?

Dynamic benchmarks solve one problem and create another.

Generators help avoid benchmark saturation by producing fresh variants of the same underlying task. But once a benchmark becomes a large parameter space, the evaluation question shifts: which regions are actually hard, and how many repeated samples are needed to prove it?

Configuration


            

Generator


            

Fresh instance

The generator is reusable, but the configuration space is the object being searched. The plots below separate the two questions: how large is the space, and where should the evaluation budget go?

Template search spaces

Each count is a set of generator settings that can produce fresh benchmark instances.

Budget allocation

Grey cells are unsampled. Outlined cells use yellow-to-red fill for observed hardness; thicker blue outlines mean more samples spent there.

Uniform sampling

Samples are spread evenly, so hard regions often get only one or two looks.

QuickScope

Revisits regions whose lower bounds stay high, building evidence.

QuickScope builds on Algorithm Configuration.

Searching a dynamic benchmark is only well-defined once you decide what you care about: raw error, error relative to complexity, robustness to a perturbation, or some other utility. QuickScope uses COUP⊕, an algorithm for optimizing user-defined utilities over large configuration spaces, then retools it for modern LLM evaluation pipelines: batched model calls, aggregate utilities, and hardness certification.

Choose the target

Set the utility: error rate, complexity-weighted error, or another scalar objective for what counts as hard.

Allocate budget

Set the model-call budget, batch size, and exploration settings; COUP uses observed utilities to revisit promising regions.

Certify hardness

With a threshold, configurations can be removed once their lower confidence bound is high enough, freeing calls for the next uncertain region.

Results explorer

Choose a benchmark, model, and utility. The plot compares the configurations found by QuickScope variants against uniform sampling where that baseline is available.

top 40