Fill available_evidence with things a reader could be pointed at: a dataset with a row count and a date range, an interview set, a screenshot, a public series you can link. Entries like "internal data" or "customer feedback" produce sections whose evidence column is equally vague, and you will not notice until drafting, when there is nothing to cite. Listing what you do not have matters as much as listing what you do; it is what stops the model from building a section around a cover-count analysis you cannot run.
The failure you will actually hit is near-equal allocations. Ask for 1600 words across five sections and a model will often return 400/350/350/300/200, which technically varies but dodges the decision. When that happens, say the budget out loud again and ask which single section it would delete if the budget dropped to 1000. The answer to that question is the real ranking, and the allocation usually fixes itself once it has been forced to name a loser.
The second failure is a number in the "What it proves" column that you never supplied. Read the output for digits, percent signs, and years, and trace each one back to your evidence list. A model asked to make a claim sound solid will reach for a plausible benchmark, and a plausible benchmark is indistinguishable from a real one at a glance. The cut list deserves the same suspicion: if the only thing cut is "Introduction" or "Conclusion", it is a strawman cut. Push back by naming a section that a ranking competitor covers and asking whether it earns its words here.
When the claims are right but the order is wrong, reorder by hand. Rerunning the whole prompt tends to regenerate claims you had already accepted, and you spend the second pass re-reviewing work you finished.