We ran the same prompt Grok's promoters brag about in 3 AI models. Only one shipped a deck you could trust.

By Deepak Sheoran, Founder and CEO, DwellFi
Grok's promoters have a favourite party trick. Type one prompt, get a finished PDF back. The way they pitch it, you describe the document you want and the model hands you something a designer would charge four figures for.
Clean. Branded. Done.
We wanted to see it hold up under a real brief, not a toy one. So we took a prompt we actually cared about and ran it through four systems the same afternoon: Grok, Claude, and our own DwellFi chat agent.
The brief was a three-day Maui trip for two. Cinematic editorial spread. Large imagery. A day-by-day itinerary. And, the part that separates a mood board from a plan, total costs. Real numbers you could budget against.
What each model actually produced
Here is the honest scorecard, because the interesting story is not that three models failed. It is that each one nailed a different single piece and dropped the rest.
.png?table=block&id=3ab9cd9a-2bee-802f-9eb6-e1a659e183a7&cache=v2)
.png?table=block&id=3ab9cd9a-2bee-80ae-b9a2-c81fef253f3b&cache=v2)
.png?table=block&id=3ab9cd9a-2bee-80c2-9fe0-c89415f38cda&cache=v2)
Grok delivered the cleanest-looking file of the three. A tidy PDF titled Hawaii_Cinematic_Travel_Plan, gold MAUI titling, elegant type, lots of white space. It looked like the promise. Then you read it. No images anywhere, despite a brief that led with imagery. And no costs. The prompt asked for a budget in plain words, and the budget never showed up. A beautiful outline is still an outline.
Claude built the most designed artifact. Dark charcoal, cream serif headings, gold accents, a real sense of layout. It even carried pricing: roughly $2,150 a guest for flights, about $1,650 a night for the hotel, near $10.2K all in for two. One problem. The numbers were labeled indicative and estimated, which is a polite way of saying invented. And the photos were not photos. Claude's own reasoning admitted it could not pull real imagery, so it dropped in teal-to-orange gradient blocks where the pictures should be. Pretty placeholders are still placeholders.
What DwellFi did differently
Same prompt. Different machine underneath.
The DwellFi agent did not treat this as one generation call. It treated it as a job with steps, and it ran them in order. First it set a goal and broke the brief into tasks. Then it pulled the skills the job needed instead of guessing.
It ran a live web search for real Maui pricing, current fares and nightly rates and activity costs, so the budget came from the open web and not from the model's imagination. It generated the imagery with a Gemini image model, real coastal and resort visuals styled to match the deck rather than clipped from anywhere. It used content-strategy to shape the narrative, so the itinerary read like a plan and not a list. It applied design-foundation with an Obsidian theme, deep and cinematic, the register the brief asked for. Then html-docs assembled the whole thing and rendered it to an eight-page PDF.
One more step the others skipped. The agent looked at its own output. A vision model read every rendered page back and checked the obvious failure modes: is the text legible, did the images load, does the budget table add up, does anything overflow. It caught its own mistakes before a human ever saw them.
The result carried what none of the others managed at once. Real web-sourced prices, not guesses: about $7,216 for two, roughly $3,608 a person, a Four Seasons night near $1,400, a Molokini snorkel run at $259 to $279 per person. Real generated images, not gradients. A complete eight-page themed deck, not an outline. And a document that had already been proofread by the system that built it.
The side-by-side
ㅤ | Grok | Claude | Gemini | DwellFi |
Real deck / PDF | Yes, clean and minimal | Yes, in-chat render | Yes, 8 pages | Yes, 8 pages |
Real images | None | Gradient placeholders | Yes, real photos | Yes, real generated |
Pricing on the page | None | Yes, but self-labeled invented | None visible | Yes, web-sourced |
Budget you can trust | No | No | No | Yes, ~$7,216 for two |
Design theme | Minimal white | Dark editorial | Cinematic dark | Obsidian, cinematic |
Self-checked before
delivery | No | No | No | Yes, vision QA |
Why this is the whole point
Prompt-to-PDF is not one problem. It is four, stacked. Get the words right. Get real numbers. Get real pictures. Wrap all of it in a design that does not embarrass you. Grok got the layout. Claude got the ambition. Each stopped at one.
The gap is not model horsepower. It is orchestration. A single model doing a single pass will always be strong somewhere and blank somewhere else. An agent that plans, pulls the right skill for each sub-job, sources live facts, generates its own assets, and grades its own work before it ships is a different category of thing.
That is the difference between an answer and a deliverable. One you read. The other you send.
Stop shipping outlines. Start shipping decks you can send. Give DwellFi the prompt that stumped Grok. See what comes back.