SquarePact Research · Legal AI Benchmark

The SquarePact Formatting BenchmarkWhat AI still breaks in Word

This report finds that every add-in tested leaves formatting defects behind on a share of its edits, and none of them signal when they have.

Full report, version 1.0

Read the full report

The complete methodology, every figure, the per-check results, and the section-by-section analysis. We'll email you the link.

Key findings

  1. Silent damage

    Between 13% and a third of the time, depending on the system, a tool edited the document and left behind a defect the user wouldn't see.

  2. Structural editing is unsolved

    Font fidelity is close to solved, with every add-in above 90%. Spacing, indentation, and the edit region itself are not, and spacing is the weakest check for every system but one.

  3. Speed doesn't predict quality

    The weakest add-in on strict pass rate was also among the fastest, and three of the four took longer on the runs they failed than on the runs they got right.

  4. No single winner

    Picking the best tool for each document reaches 94.2%. The best single tool reaches 80.2%, so 14 points of headroom sit in choosing correctly rather than in any one product.

  5. Cost and latency, per configuration

    Dollars, tokens, and seconds per edit are reported for every system tested, including the two API-key modes, where a frontier model at high reasoning cost the most and edited no more cleanly.

The damage rate

The damage rate counts the portions where a system made the edit and left a defect behind. It's the failure a reviewer is least likely to catch, because the document comes back looking finished.

Share of portions edited and damaged

Lower is better. Counts a portion when check 1 passed but at least one other check failed.

Best Other systems
SquarePact
12.8% 11 of 86
claude-api-in-
squarepact*
16.7% 5 of 30
Claude for Word
20.9% 18 of 86
openai-api-in-
squarepact*
23.3% 7 of 30
Microsoft Copilot
31.4% 27 of 86
Grok for Word
32.6% 28 of 86
Figure 4. Paired McNemar on the same portions: SquarePact vs Copilot p = 0.0015 and vs Grok for Word p = 0.0023, both significant. Against Claude for Word the margin is 12 portions to 5, p = 0.14, which this sample cannot separate. Rows marked * are our own add-in on a customer API key, scored on 30 portions rather than 86; they look safe here only because they decline so many edits.

Who built this

SquarePact and this report are built and maintaned by Actualization.AI, a Tampa, Florida company that spun out of the Advancing Machine and Human Reasoning Lab at the University of South Florida. Our CEO, Professor John Licato, directs that lab and has worked in AI for more than twenty years, with over 100 peer-reviewed publications and NSF SBIR federal research funding in language models and document reasoning. Our VP of Engineering, Dr. Animesh Nighojkar, came out of the same lab.

We've spent years studying why language models fail at professional document work rather than noting that they do, and this benchmark existed internally as our own development metric long before it existed publicly. We're a participant in the comparison rather than solely a neutral referee, so the full dataset and the analysis script ship alongside the report and anyone can re-run it. More info about us here.

John Licato, Ph.D.
John Licato, Ph.D.CEO, Actualization AI. Director, AMHR Lab, University of South Florida
Animesh Nighojkar, Ph.D.
Animesh Nighojkar, Ph.D.VP of Engineering, Actualization AI
The SquarePact team presenting on agentic AI The SquarePact team at Embarc Collective The SquarePact team with event panelists

Open data

The dataset and analysis script are public. Every number in the report can be recomputed from the released CSV.

License
Permissive. Free to use for research or product work; attribution and a link back to this page are required.
Inquiries
john@actualization.ai. If you've built a tool you'd like measured against this set, we'd like to hear from you.

Citation

@techreport{squarepact_formatting_benchmark_2026,
  title  = {The SquarePact Formatting Benchmark},
  author = {Licato, John and Nighojkar, Animesh and Vaidya, Darsh
            and Sajithkumar, Govind},
  year   = {2026},
  note   = {Version 1.0},
  institution = {Actualization AI, Inc.},
  url    = {https://squarepact.com/formattingBench/}
}

Results are tied to a version and a date. The set is updated on a regular cadence, so scores from different versions aren't a like-for-like comparison.

Full report, version 1.0

Read the full report

We use your email to send the report and occasional research from the same team. No sharing, unsubscribe any time. If the form doesn't load, email john@actualization.ai.