I have sat on more source-selection panels than I care to count, and the last several have all had the same strange quality. Nothing is obviously bad. The weak bidder no longer submits the response with the mismatched fonts, the boilerplate from a different client, and the paragraph that answers a question we didn't ask. Everybody's submission is tight. Everybody's approach section maps cleanly to our requirements. Everybody's risk register is populated and plausible. And the panel, which has always used the written response to thin the field, finds it has nothing to thin with.

This is not a hunch about vibes. Loopio's annual survey of proposal teams tracked generative AI use going from 34% of teams in 2023 to 79% in the most recent report, with 84% of adopters using it at least weekly. Whatever the exact figure is in your market, the direction is not in dispute: the written proposal is now a machine-assisted artifact on essentially every desk it lands on, including the desks of the firms you would not hire.

I want to be careful here, because the reflex is to treat this as vendors cheating. It isn't. If I were bidding, I would use every tool available to produce a clear, compliant, well-organised response, and I would think less of a firm that didn't. The problem is not on the vendor's side of the table. It's on mine. My evaluation method quietly assumed the document was evidence of something, and nobody went back and checked whether that assumption still holds.

What the proposal was actually measuring

Scoring a written response was never really about writing. The document was a proxy, and it stood in for three things I genuinely needed to know.

First, that somebody senior had read our requirements closely enough to be specific about them — because being specific is expensive, and generic is what you write when nobody with authority spent time on it. Second, that the firm could organise itself against a deadline, since a bid is a small program with a fixed date, a dependency chain, and multiple contributors who have to be made to agree. Third, that they had the bench to spare. A serious response cost a serious firm real senior hours it could have billed elsewhere, and choosing to spend them on us was information.

Michael Spence won a Nobel for formalising the mechanism underneath all three: a signal only carries information when it's costly, and — this is the part people forget — when it's more costly for the candidates who lack the quality being signalled. That's the whole trick. The reason a hard thing tells you something is that it is disproportionately hard for the wrong people. Generative AI didn't lower the cost of a polished proposal a little. It lowered it toward zero, and it lowered it furthest for the bidder who previously couldn't clear the bar. The signal didn't get weaker. It inverted, and now rewards whoever configured their tooling best.

Three ways evaluation panels are still getting this wrong

1. The rubric still scores the artifact

Open your last scoring sheet and count the criteria that are satisfied by the document itself: clarity of approach, quality of written response, demonstrated understanding of requirements. Those were always shorthand for qualities of the team, and the document was the only window we had into them. The window is now a painting. Leaving those weights untouched means a meaningful share of the award is being decided by tool access.

2. We test for the presence of words, not the presence of people

Named key personnel appear in the response with tailored bios explaining why they're perfect for this engagement. Then delivery starts and a different bench shows up — not out of malice, usually, but because the named people were never actually committed, and nothing in the evaluation forced them to be. A résumé in a PDF is now the cheapest object in the transaction.

3. We treat the proposal as the contract's origin story

Panels lean on the written response as the record of what was promised, which matters enormously later, when something slips and somebody goes looking for what was agreed. But if the response was generated, the commitments in it may never have been argued over inside the bidding firm at all — the same failure I've described in plans nobody agreed to, one organisation upstream. A commitment nobody debated is not a commitment. It's formatting.

The principle

A proposal is no longer evidence of capability. It's evidence of access to the same tools everyone else has. Weight your evaluation toward the things a vendor cannot generate — their actual people, answering unscripted questions, about your actual problem.

Evaluate the team, not the document

Here's what I've changed, and none of it is exotic. The written response still gates compliance — it has to be responsive, it has to be complete — but its share of the score comes down, and that weight moves to a live working session. Not a pitch: a working session, with the named delivery lead and the named architect in the room, on a real problem from our program that they haven't seen, where the panel gets to ask follow-ups. Prose cannot survive a third follow-up question. People either know the domain or they visibly don't, and twenty minutes of that tells me more than sixty pages ever did.

Alongside it, three smaller changes. I ask every bidder to tell me about a program of this shape that went badly and what they changed afterward — a question with no good generated answer, because the specificity that makes it credible only comes from having lived it. I make key-personnel commitments contractual, with named individuals, a minimum allocation, and a substitution clause that requires our consent, so the bios in the PDF have to become people on the calendar. And I ask for disclosure of how AI was used in the response, not to penalise it, but because a firm that can tell me which human signed off on the content is demonstrating the exact governance discipline I'm going to need from them on delivery. The ones who can answer that cleanly are telling me something real. So are the ones who can't.

How to actually do this
  • Reweight the rubric. Anything scored purely from the document is now measuring tool access — move that weight to live, observed evaluation.
  • Run a working session, not a pitch. Named delivery lead, an unseen problem from your program, and the right to ask follow-ups.
  • Ask what went wrong last time. Specific, unflattering, verifiable detail is the signal that's still expensive to fake.
  • Make key personnel contractual — named, allocated, and not substitutable without your consent.
  • Ask how AI was used and who reviewed the output. You're not policing the tool; you're auditing their governance before you're depending on it.

The bottom line

The Army taught me to evaluate a unit by watching it operate, not by reading what it said about itself, and that instinct has never been more relevant to commercial procurement than it is right now. We built our selection process in an era when producing a competent document was hard, and we let the document stand in for competence because the correlation was good enough to run on. That correlation has broken. Pretending otherwise doesn't make evaluation neutral — it makes it random, and random selection on a program you're about to hand nine figures of scope to is a risk nobody put in the register. Read the proposal for compliance. Then put the people in a room and find out who you're actually buying.