The test in one minute

On September 16, 2026, AI Stack Route gave the logged-out ChatGPT web experience one synthetic small-business writing task three times. Each submission went into a separate desktop browser tab and used the exact same prompt. We did not use an account, upload files, add follow-up instructions, or provide personal, customer, or confidential business information.

The fictional assignment asked for a 120–150-word homepage section and three email subject lines under 45 characters for a small bookkeeping service. The prompt supplied four fact bullets, named the audience and tone, and prohibited testimonials, guarantees, certifications, locations, tenure, results, and added services.

The interface did not disclose the model or version. These three runs form one limited test of one task under one access condition. They do not establish an overall ChatGPT rating or show how another model, paid plan, product, or task would perform.

What we asked ChatGPT to create

The prompt was designed to test two different kinds of instruction at once. First, the draft had to preserve a fixed set of business facts without filling in missing information. Second, it had to meet measurable format constraints: a homepage section within a word range and three subject lines below a character limit.

That combination matters for customer-facing copy. A draft can sound polished while still missing the requested length, repeating a call to action, or adding a claim the business cannot support. Publishing the exact prompt makes the test reproducible and lets readers judge the outputs against the same brief.

Read the exact synthetic prompt
You are helping a fictional three-person bookkeeping service called Juniper Ledger. Using only the facts below, draft:

1. A homepage section of 120–150 words with a plain-English headline, a short body, and a call to action.
2. Three email subject lines, each under 45 characters.

Audience: freelancers and local service businesses.

Facts:
- Monthly bookkeeping plans start at $299.
- A free 20-minute consultation is available.
- Service is remote.
- Clients receive monthly categorized transaction reports and a quarterly check-in.

Tone: clear, calm, and practical.

Do not invent testimonials, guarantees, certifications, locations, years in business, results, or additional services. If information is missing, do not fill it in. Return only the homepage section and the three subject lines.

What stayed consistent across all three runs

Each draft retained every supplied service fact: plans starting at $299 per month, a free 20-minute consultation, remote service, monthly categorized transaction reports, and a quarterly check-in. We did not find an invented testimonial, guarantee, certification, location, tenure claim, result, or added service in the three outputs.

Every subject line stayed below the requested 45-character limit. The language also remained clear, calm, and practical. Those observations support a narrow conclusion: for this task, the logged-out experience produced coherent first drafts grounded in most of the explicit instructions.

  • Run 1 subject lines: 33, 31, 39 characters.
  • Run 2 subject lines: 34, 32, 39 characters.
  • Run 3 subject lines: 36, 36, 39 characters.

Where the drafts missed the brief

The largest consistent miss was length. Counting the headline, body, and call to action using the recorded measurement rule, the three homepage sections contained 58, 73, and 95 words. None reached the requested 120–150-word range.

Run two also repeated the consultation call to action in consecutive lines. Run three added the phrase “No complicated process.” The supplied facts did not describe the process, so that phrase would need to be removed or supported before publication. Calling the phrase unsupported does not mean the service is complicated; it means the test prompt supplied no evidence for the claim.

A fluent draft can therefore look finished while still failing a check that is easy to measure. For a business owner, the useful question is not whether the first draft sounds confident. It is whether every fact, constraint, and promise survives review.

  • Run 1: 58 words against the requested 120–150.
  • Run 2: 73 words against the requested 120–150.
  • Run 3: 95 words against the requested 120–150.

What this result means for small-business copy

For this narrow task, ChatGPT was useful as a drafting surface. It turned structured facts into readable copy and produced subject-line options within the stated character limit. It was not ready for unattended publication because all three drafts missed the word-count requirement and one introduced an unsupported process claim.

That distinction is more useful than a broad verdict. A shorter section might be appropriate in a different layout, but the assignment explicitly required 120–150 words. The failure is about instruction compliance in this test, not a universal rule about writing quality.

The practical role is first draft, followed by a human check. The reviewer should have the approved source facts and the original brief open beside the output. If a statement cannot be traced to those facts or another approved source, remove it or verify it before the copy reaches a customer.

A five-point review before publishing AI-written business copy

Use the same review order each time so polished wording does not distract from missing evidence. Start with objective checks, then edit for voice and usefulness.

  • Check every fact against an approved source. Confirm prices, offer details, deliverables, availability, and qualifications instead of relying on the draft.
  • Count words and characters. Verify each stated length, number of items, heading requirement, and format instruction with a separate check.
  • Remove or substantiate benefits and process claims. Watch for easy, simple, seamless, guaranteed, faster, better, and similar language that may need evidence.
  • Check repetition, audience fit, and brand voice. Remove duplicate calls to action and make sure the copy speaks to the intended reader without invented pain points.
  • Confirm the final action and offer details. The button, consultation, price, destination, and next step should match what the business can actually provide.

What this test does not establish

The test did not compare ChatGPT with another product and does not support a ranking. It did not test paid plans, account features, memory, file uploads, research, privacy controls, images, analysis, or any task beyond the single fictional writing brief.

The three outputs are repeated runs inside one evaluation, not three independent studies. The timestamps document when the outputs were captured; they are not a speed benchmark. Because the logged-out interface did not name the model or version, the results should not be assigned to a specific model.

  • No numeric quality score or best-writer claim.
  • No claim that ChatGPT always preserves facts or always misses constraints.
  • No estimate of time saved, revenue gained, or conversion improvement.
  • No conclusion about sensitive-data use or privacy settings.

A practical way to run your own low-risk test

Choose one repeatable, low-risk draft and use non-sensitive information. Write down the approved facts, prohibited claims, audience, format, and pass-or-fail constraints before opening the tool. Run the same task more than once, then compare the results with the original brief rather than choosing the most polished version by instinct.

Record what changed, what stayed accurate, and how much correction each result needed. Publish only after a person accountable for the business has approved the facts, claims, offer, voice, and final call to action. For current product and plan information, use the official OpenAI sources listed below instead of inferring features from this test.

Sources reviewed

First-hand evidence and current official pages checked for this article. Product details can change after the checked date.