This report finds that every add-in tested leaves formatting defects behind on a share of its edits, and none of them signal when they have.
Full report, version 1.0
The complete methodology, every figure, the per-check results, and the section-by-section analysis. We'll email you the link.
Between 13% and a third of the time, depending on the system, a tool edited the document and left behind a defect the user wouldn't see.
Font fidelity is close to solved, with every add-in above 90%. Spacing, indentation, and the edit region itself are not, and spacing is the weakest check for every system but one.
The weakest add-in on strict pass rate was also among the fastest, and three of the four took longer on the runs they failed than on the runs they got right.
Picking the best tool for each document reaches 94.2%. The best single tool reaches 80.2%, so 14 points of headroom sit in choosing correctly rather than in any one product.
Dollars, tokens, and seconds per edit are reported for every system tested, including the two API-key modes, where a frontier model at high reasoning cost the most and edited no more cleanly.
The damage rate counts the portions where a system made the edit and left a defect behind. It's the failure a reviewer is least likely to catch, because the document comes back looking finished.
Share of portions edited and damaged
Lower is better. Counts a portion when check 1 passed but at least one other check failed.
SquarePact and this report are built and maintaned by Actualization.AI, a Tampa, Florida company that spun out of the Advancing Machine and Human Reasoning Lab at the University of South Florida. Our CEO, Professor John Licato, directs that lab and has worked in AI for more than twenty years, with over 100 peer-reviewed publications and NSF SBIR federal research funding in language models and document reasoning. Our VP of Engineering, Dr. Animesh Nighojkar, came out of the same lab.
We've spent years studying why language models fail at professional document work rather than noting that they do, and this benchmark existed internally as our own development metric long before it existed publicly. We're a participant in the comparison rather than solely a neutral referee, so the full dataset and the analysis script ship alongside the report and anyone can re-run it. More info about us here.
The dataset and analysis script are public. Every number in the report can be recomputed from the released CSV.
Citation
@techreport{squarepact_formatting_benchmark_2026,
title = {The SquarePact Formatting Benchmark},
author = {Licato, John and Nighojkar, Animesh and Vaidya, Darsh
and Sajithkumar, Govind},
year = {2026},
note = {Version 1.0},
institution = {Actualization AI, Inc.},
url = {https://squarepact.com/formattingBench/}
}Results are tied to a version and a date. The set is updated on a regular cadence, so scores from different versions aren't a like-for-like comparison.
Full report, version 1.0
We use your email to send the report and occasional research from the same team. No sharing, unsubscribe any time. If the form doesn't load, email john@actualization.ai.