How2Everything: Mining the Web for How-to Procedures to Evaluate and Improve LLMs

Yapei Chang, Kyle Lo, Mohit Iyyer, Luca Soldaini
Proceedings of the 43rd International Conference on Machine Learning, PMLR 306:12779-12825, 2026.

Abstract

Generating step-by-step "how-to" procedures is a key LLM capability: how-to advice is commonly requested in chatbots, and step-by-step planning is critical for complex reasoning tasks. Yet, measuring and improving procedural validity at scale on real-world tasks remains challenging and understudied. We introduce How2Everything, a scalable framework to evaluate and improve goal-conditioned procedure generation. Our pipeline How2Mine extracts and rewrites 351K procedures from 980K web pages across 14 topics, and can scale to larger corpora. From this pool we build How2Bench, a 7K-example evaluation set balanced across topics. We also introduce How2Score, an evaluation protocol that uses an LLM judge to detect whether a generation contains any critical failure that would prevent achieving the goal. For low-cost, reproducible evaluation, we distill a frontier judge into an open 8B model achieving 80.5% agreement with human annotators. How2Bench reveals clear scaling trends across model size and training stages, providing signal early in pretraining. Finally, RL using How2Score as a reward improves performance on How2Bench by $>$10 points across three base models without systematic regressions on standard benchmarks, with gains not primarily explained by source-document memorization or superficial format compliance. We release all code and data at https://github.com/lilakk/how2everything.

Cite this Paper


BibTeX
@InProceedings{pmlr-v306-chang26e, title = {{H}ow2{E}verything: Mining the Web for How-to Procedures to Evaluate and Improve {LLM}s}, author = {Chang, Yapei and Lo, Kyle and Iyyer, Mohit and Soldaini, Luca}, booktitle = {Proceedings of the 43rd International Conference on Machine Learning}, pages = {12779--12825}, year = {2026}, editor = {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto}, volume = {306}, series = {Proceedings of Machine Learning Research}, month = {06--11 Jul}, publisher = {PMLR}, pdf = {https://raw.githubusercontent.com/mlresearch/v306/main/assets/chang26e/chang26e.pdf}, url = {https://proceedings.mlr.press/v306/chang26e.html}, abstract = {Generating step-by-step "how-to" procedures is a key LLM capability: how-to advice is commonly requested in chatbots, and step-by-step planning is critical for complex reasoning tasks. Yet, measuring and improving procedural validity at scale on real-world tasks remains challenging and understudied. We introduce How2Everything, a scalable framework to evaluate and improve goal-conditioned procedure generation. Our pipeline How2Mine extracts and rewrites 351K procedures from 980K web pages across 14 topics, and can scale to larger corpora. From this pool we build How2Bench, a 7K-example evaluation set balanced across topics. We also introduce How2Score, an evaluation protocol that uses an LLM judge to detect whether a generation contains any critical failure that would prevent achieving the goal. For low-cost, reproducible evaluation, we distill a frontier judge into an open 8B model achieving 80.5% agreement with human annotators. How2Bench reveals clear scaling trends across model size and training stages, providing signal early in pretraining. Finally, RL using How2Score as a reward improves performance on How2Bench by $>$10 points across three base models without systematic regressions on standard benchmarks, with gains not primarily explained by source-document memorization or superficial format compliance. We release all code and data at https://github.com/lilakk/how2everything.} }
Endnote
%0 Conference Paper %T How2Everything: Mining the Web for How-to Procedures to Evaluate and Improve LLMs %A Yapei Chang %A Kyle Lo %A Mohit Iyyer %A Luca Soldaini %B Proceedings of the 43rd International Conference on Machine Learning %C Proceedings of Machine Learning Research %D 2026 %E Tong Zhang %E Miroslav Dudik %E Martin Jaggi %E Alekh Agarwal %E Sharon Li %E Dale Schuurmans %E Jerry Zhu %E Felix Berkenkamp %E Hanze Dong %E Alberto Bietti %F pmlr-v306-chang26e %I PMLR %P 12779--12825 %U https://proceedings.mlr.press/v306/chang26e.html %V 306 %X Generating step-by-step "how-to" procedures is a key LLM capability: how-to advice is commonly requested in chatbots, and step-by-step planning is critical for complex reasoning tasks. Yet, measuring and improving procedural validity at scale on real-world tasks remains challenging and understudied. We introduce How2Everything, a scalable framework to evaluate and improve goal-conditioned procedure generation. Our pipeline How2Mine extracts and rewrites 351K procedures from 980K web pages across 14 topics, and can scale to larger corpora. From this pool we build How2Bench, a 7K-example evaluation set balanced across topics. We also introduce How2Score, an evaluation protocol that uses an LLM judge to detect whether a generation contains any critical failure that would prevent achieving the goal. For low-cost, reproducible evaluation, we distill a frontier judge into an open 8B model achieving 80.5% agreement with human annotators. How2Bench reveals clear scaling trends across model size and training stages, providing signal early in pretraining. Finally, RL using How2Score as a reward improves performance on How2Bench by $>$10 points across three base models without systematic regressions on standard benchmarks, with gains not primarily explained by source-document memorization or superficial format compliance. We release all code and data at https://github.com/lilakk/how2everything.
APA
Chang, Y., Lo, K., Iyyer, M. & Soldaini, L.. (2026). How2Everything: Mining the Web for How-to Procedures to Evaluate and Improve LLMs. Proceedings of the 43rd International Conference on Machine Learning, in Proceedings of Machine Learning Research 306:12779-12825 Available from https://proceedings.mlr.press/v306/chang26e.html.

Related Material