How We Test Whether an AI Draft Is Actually Good

TT

Tripdash Team

July 7, 2026

An AI draft that looks polished isn't the same as one that's good. Here's the internal bar Tripdash holds every draft to before it reaches an advisor or a traveler.

6 minute read
A travel advisor reviewing an AI-generated itinerary draft on a laptop alongside handwritten notes

A well-formatted itinerary is easy to produce. Ask almost any AI model to plan five days in Lisbon and it will hand back something that reads smoothly: a morning walk through Alfama, a sunset at a viewpoint, a restaurant recommendation with a nice description. It looks like a finished trip. Whether it actually works as a trip is a different question, and it's the one we spend most of our internal effort answering before any AI-assisted feature reaches an advisor or a traveler.

This matters more in travel than in most categories where AI drafting gets used. A draft blog post that reads slightly off can be edited in a few minutes. A draft itinerary that's slightly off might send someone across a city and back during the two hours a museum happens to be closed, or assume a rental car is available on a Sunday when the local agency doesn't open until Monday. The failure isn't stylistic. It's operational, and it shows up when someone is standing in an airport, not reading a draft on a screen.

Before describing what a good plan looks like, it helps to explain how we decide whether one actually is good.

The gap between fluent and correct

Language models are trained to produce text that sounds right. That's genuinely useful for turning a pile of preferences into a readable itinerary, but fluency and correctness are not the same property, and it's easy to mistake one for the other when reading quickly. A draft can use the right neighborhood names, the right tone, the right structure, and still fail on the details that determine whether the day actually works: whether that restaurant takes walk-ins on the night in question, whether that hike is reasonable for someone who mentioned a knee injury, whether the pacing leaves room to eat lunch instead of rushing between three timed entries.

The uncomfortable part is that a fluent-but-wrong draft is often more dangerous than an obviously wrong one. A glaring error tends to get caught immediately. A plan that's mostly right and confidently wrong on one constraint is the one that slips through if the only check applied is "does this read well."

That's why a draft is never treated as a finished product here. It's treated as a starting point that has to earn its way to a traveler, and it earns that by clearing a review built to catch something other than polish.

Human advisor review as the actual gate

Every AI-assisted draft that reaches a traveler has passed through a human advisor reading it as a professional, not as a proofreader. That distinction matters. A proofreader checks whether the sentences are clean. An advisor checks whether the trip is real: whether the timing between activities makes sense given actual transit conditions, whether the recommendations fit the traveler's stated budget rather than a plausible-sounding one, whether a pace that looks fine on paper would exhaust a family with young kids or feel rushed for someone who asked for a slower week.

This part of the process isn't automated, and there's no plan to automate it. The AI is good at generating options and structure quickly. It isn't accountable for the trip going well, and it isn't treated as though it were. The advisor is the one who has actually planned trips like this before, who knows a "10-minute walk" on a map can be a 25-minute walk uphill in July heat, and who is willing to send a draft back rather than let it go out because it looked finished.

Review isn't a symbolic step where someone skims a wall of text hoping nothing's wrong. Unstructured review tends to catch whatever's visually obvious and miss whatever's quietly wrong underneath, which is usually the more expensive mistake. So the review itself checks specific things, in a specific order, rather than relying on a general impression of quality.

Checking against constraints that actually exist

The biggest source of bad AI drafts isn't bad writing, it's a draft built on assumptions instead of facts. A model asked to plan a trip will often fill gaps with plausible defaults: assuming a museum is open, assuming a restaurant has room for a group, assuming a transfer takes as long as a map suggests rather than as long as it actually takes at that hour, in that traffic, with that transit schedule. Plausible is not the same as true, and travel is a domain where the two diverge constantly.

So a meaningful part of quality checking is comparing the draft against constraints that exist independently of the AI's output: actual opening hours and seasonal closures, actual availability for the dates in question, actual pricing rather than a general estimate, actual transit times rather than straight-line distances. None of this is exotic. It's the same due diligence any competent advisor would do by hand, applied deliberately, because a draft looks equally confident whether its underlying assumption was correct or not.

Budget fit gets checked the same way. A plan can hit every preference a traveler listed and still be unusable if the total cost drifts past what they wanted to spend, or if it assumes a room type or flight class that isn't actually available at that price point. An AI draft doesn't reliably know when it's doing this. A human checking against the traveler's actual budget does.

Pacing is a quality dimension, not a nice-to-have

It's tempting to treat pacing as a soft, subjective concern compared to something concrete like availability. In practice it's one of the more common reasons a technically correct itinerary still isn't a good one. A day can have every venue open, every reservation valid, every price accurate, and still be a bad day if it has someone doing a four-hour hike in the morning followed by a late dinner across town with no buffer in between, or if it packs six attractions into an afternoon because each one individually seemed reasonable to include.

Pacing failures are subtle because they don't show up as a factual error. They show up as a trip that reads fine and doesn't work when someone tries to actually live it. Checking for this means asking questions a model doesn't automatically prioritize: is there room to be tired, is there room for a meal that isn't rushed, does a family with kids get the same day a solo traveler would, does a trip advertised as relaxed actually have any slack in it.

What we check before a draft ships

Quality dimensionWhat "good" looks likeHow it's checked
AvailabilityVenues, restaurants, and transport are actually bookable or open on the dates proposedCross-checked against current hours, seasonal closures, and booking status rather than general knowledge
PacingDays leave realistic time for transit, meals, and rest given the traveler's stated pace preferenceAdvisor reads the day as a lived sequence, not a checklist, and flags anything overpacked or oddly spaced
Budget fitTotal cost matches the traveler's stated range, using real pricing for the dates in questionLine-item comparison against the stated budget, not an average or estimate
Constraint accuracyPhysical, dietary, and accessibility needs mentioned by the traveler are actually reflected in the planAdvisor checks each stated constraint against each relevant recommendation individually
Local plausibilityRecommendations reflect how a place actually works, not a generic version of itAdvisor judgment, informed by direct destination experience or a trusted local source

How a failed review changes the product

A draft that fails review doesn't just get corrected and forgotten. When the same kind of issue keeps showing up, that's a signal about the underlying feature, not just one itinerary. If advisors keep catching the same pacing problem in a particular type of trip, or the same wrong assumption about a category of venue, that pattern gets fed back into how the drafting step works and what it's asked to check for itself before a human ever sees it.

This is slower than shipping a feature the moment it produces plausible-looking output, and that's intentional. The goal isn't a demo that looks impressive once. It's a drafting step that keeps getting more useful to advisors because it keeps making fewer of the mistakes review has to catch. That only happens if failures are actually tracked and fed back in, instead of being fixed once and forgotten.

The bar, in plain terms

This is the bar we hold internally, and it's the reason an AI-assisted draft from Tripdash goes through an advisor rather than straight to a traveler. The technology makes the first version faster to produce. Whether that version is actually good is still a human judgment, checked against the real world, every time.

Frequently asked questions

No. Every AI-assisted draft passes through a human advisor before it reaches a traveler. The AI helps produce a fast first version, but the advisor is the one checking whether the trip actually works.

Mostly assumption-based errors: a venue assumed to be open when it isn't, a transfer time that looks fine on a map but ignores real traffic or transit schedules, pacing that packs too much into a day, or a budget that quietly drifts past what the traveler asked for.

An advisor reads the day as a sequence someone would actually live through, not as a checklist of open venues and valid reservations. That means checking whether there's realistic time for transit, meals, and rest given the traveler's stated pace preference.

It gets corrected before anything reaches the traveler, and if the same kind of issue keeps showing up across drafts, that pattern gets fed back into how the drafting step works so the same mistake is less likely next time.

Because a draft can be fluent and still be wrong on details that matter operationally, like availability or budget fit. Those errors don't look like errors on the page, so an advisor who has actually planned trips is the one accountable for catching them.

More from Tripdash