Abstract
Policies with similar mean returns can differ sharply in rare failures, yet estimating lower-tail conditional value-at-risk (CVaR) accurately can require many costly rollouts. When different conditional components of a stochastic workflow can be queried separately, we ask how to allocate a fixed evaluation budget to estimate a fixed policy's CVaR most accurately. We derive a tail influence for each queryable conditional law that aggregates how its uncertainty affects CVaR across every Bellman reuse. Its variance yields the fixed-design efficiency bound and the oracle Neyman allocation. Tail-Influence Sampling (TIS) estimates these influence scales from a pilot model and reallocates fresh queries toward kernels that matter most for the tail; a visitation-anchored variant protects against pilot underallocation. Under fixed dimension and a positive quantile margin, TIS attains oracle asymptotic variance and first-order MSE including pilot cost, while the anchored variant is within a factor two of the oracle. We also characterize an exact-grid regime in which tail- and mean-optimal allocations coincide. On CliffWalking, TIS reduces MSE by 41% versus learned occupancy and 76% versus complete rollouts at the same charged transition budget. In frozen language-model review workflows, anchored TIS beats an equally regularized mean-influence blend in 23 of 24 MMLU-Pro settings and reaches 2.4-3.4times lower MSE than rollouts on six-call FinQA reviews.
Community
We introduce Tail-Influence Sampling (TIS) to estimate performance in the worst runs more accurately with a fixed evaluation budget.
When workflow components can be sampled separately, TIS learns which ones need more queries. On CliffWalking, it achieves 76% lower estimation error (MSE) than complete rollouts at the same query budget, including pilot costs. We also test it on LLM review workflows.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Speculative Evaluation of Stochastic LLMs (2026)
- Information Limits of Multistage Inventory Control: Learning, Valuation, and Censoring (2026)
- Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation (2026)
- Bellman-Certified Rounding for Sparse Policy Deployment in MDPs (2026)
- Exponential Hardness of Off-Policy Evaluation under History-Dependent Logging (2026)
- Target-Aware Sequential Inference: Pooled versus Stratified Anytime-Valid Designs (2026)
- Evaluator Ensembles Under Reward Hacking: Covariance Geometry and Finite-Search Guarantees (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.38096 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 1
Collections including this paper 0
No Collection including this paper