← Back to the journal

Design a Concierge Pilot That Tests One Customer Outcome

Scope a SaaS concierge pilot with one customer outcome, a support budget, measurable success criteria, and a clear decision about paid continuation.

Founder and customer review a concierge pilot: support tickets become an approved weekly report, while the pilot brief flags manual effort for scope revision and leaves paid continuation undecided.
Conceptual editorial artwork · Generated with AI for FindVex

Before inviting customers into a concierge pilot, agree on the result they need, how much help you will provide, and the terms for continuing. Otherwise, a useful result can leave you unsure whether the customer values the software, your personal attention, or both.

A concierge pilot is a limited engagement in which your team manually performs or supports parts of a proposed software workflow. The customer knows people are involved. You might prepare inputs, check AI output, or deliver the result yourself.

The brief below helps you test one customer outcome while tracking the work required to deliver it. Adapt its boundaries to your workflow; the example’s numbers are planning assumptions, not benchmarks.

Choose a task the customer already needs to finish

Ask a prospective participant to walk through the last time they completed the task. Find out what triggered it, who did the work, which tools they used, and what made the result acceptable. Then ask when the task must happen again, where delays or rework become costly, and who can approve spending to improve the process.

Starting with the current alternative follows April Dunford’s positioning method: identify what customers would do without your solution, then connect your different capabilities to value and the customers who care about it. Dunford’s positioning guide explains those connections.

Record the customer’s wording separately from your interpretation. If someone describes difficulty reconciling conflicting notes, don’t silently turn that into a request for an AI dashboard. Use FindVex’s pain-point evidence sheet to organize the observations, then confirm the problem with prospective participants.

Recruit people with the same recurring task and constraints. Enthusiasm for testing software is useful, but participants also need a reason to complete the work you are testing.

Manual delivery can help you learn what to build. Paul Graham recommends doing work by hand early and gradually automating bottlenecks. His discussion of manual delivery provides a rationale for that approach; a pilot still needs its own evidence of customer value and delivery cost.

Write the outcome and its acceptance rule together

“Use the assistant” describes an activity. “Produce a weekly report the operations lead can approve” describes a result, but still needs an acceptance rule.

Write down the baseline, target, measurement method, and deadline before delivery starts. Strategyzer’s Test Card separates the hypothesis, test, metric, and success threshold. See the original Test Card explanation.

Add a quality requirement alongside any speed target. Decide who reviews the output and what counts as a material correction. Include the customer’s preparation, review, and correction time so that a faster first draft cannot hide extra work later.

For a repeated workflow, specify how many deliveries must meet the rule. Keep each cycle’s result visible, even if you also calculate an average.

Question Evidence to collect
Did the customer benefit? Accepted output and comparison with the baseline
What produced the benefit? Software steps and manual interventions
Can the team deliver it economically? Delivery time, support time, and direct costs
Will the buyer continue? A decision on a specific paid scope, with payment status recorded separately

One successful report cannot answer all four questions. A small pilot can justify a further test, but cannot establish retention or demand across a market.

Put a boundary around your help

Write a scope both sides can understand: participant, required inputs, deliverable, duration, and customer responsibilities. Name the adjacent work you will exclude.

Choose a duration that includes a meaningful work cycle and, where practical, a repeat. A weekly workflow needs a different observation window from a quarterly planning process. Ending before the customer’s next real task leaves the repeat-use question unanswered.

Budget for onboarding, data preparation, output review, corrections, and follow-up. Track setup separately from recurring delivery. This lets you see whether the work becomes easier and which costs continue with every customer cycle.

For an AI workflow, log what the model produced and what a person changed. A polished final answer can hide extensive manual rewriting. Disclose who will review customer inputs and outputs, and agree on approved data before requesting it.

When a participant requests extra work, record the request and its reason. Keep it outside the pilot or explicitly revise the scope and measurement. Mark results under a revised scope separately so you can see which agreement produced them.

Worked example: a weekly support report

Imagine a two-person startup building an AI tool that turns support tickets into a weekly issue report. It recruits two small SaaS support teams that already prepare such a report and can provide an approved ticket export. Every number and result in this example is hypothetical.

Each customer measures two reporting cycles before the pilot. Assume each team spends 90 minutes per report, including preparation and review.

The proposed pilot has these terms:

Item Agreed scope
Duration Three weekly reporting cycles
Input Up to 200 tickets per team per cycle, in a fixed export format
Deliverable An issue report with links to supporting tickets, reviewed by the support lead
Customer outcome Each of the three reports approved with no unsupported priority claims and no more than 30 minutes of customer effort per report
Time measurement Export preparation, review, and corrections included; quality checked against the same rubric used for the baseline
Founder effort Up to two setup hours per account and 45 minutes per weekly delivery, including questions and corrections
Exclusions Live integrations, customer replies, and product-roadmap recommendations
Proposed continuation $300 per month for four reports under the same volume and service limits, with founder review disclosed

The pilot is free in this example. It tests delivery while leaving payment unresolved until the buyer makes a commercial decision.

Suppose both teams meet the acceptance rule on all three reports, but the founder spends 75 minutes on every report. At an assumed internal labor cost of $60 per hour, recurring labor for one account would be:

4 reports × 75 minutes ÷ 60 × $60 = $300 per month.

That consumes the entire proposed monthly price before compute, setup, or overhead. The customers received the agreed result, but delivery exceeded the effort cap. Identify which manual step caused the overrun and test whether you can reduce it without weakening the result before committing to the proposed ongoing scope.

Now suppose one buyer agrees to continue and the other declines. Record those decisions and their reasons, then track whether the accepting buyer actually pays. Check whether that buyer expects the same human review. Even a completed purchase would establish willingness to pay for this managed service, leaving demand for a future self-service product untested.

Agree on the commercial decision before kickoff

Show the proposed continuation scope and price before the pilot begins. Name the decision-maker and book a closing review so the customer has a concrete offer to evaluate throughout the work.

A paid pilot tests willingness to pay for the pilot package, including its human assistance. A free pilot can investigate delivery, but leaves payment unresolved. Neither arrangement proves renewal.

If you charge, specify what the fee covers and when it is due. For a free pilot, ask for the commitments needed to run it: timely inputs, a named reviewer, and attendance at the closing review.

At that review, compare the results with the original criteria:

Result Next decision
Outcome met, effort acceptable, buyer accepts Confirm the paid scope and payment terms, then track repeat use
Outcome met, effort too high Test a narrower workflow, reduce a specific manual step, or evaluate a price the buyer will accept
Outcome missed Investigate inputs, output quality, onboarding, and whether the problem matters enough
Buyer declines despite a useful outcome Record the reason and decide which assumption, if any, deserves another test

A procurement delay is an unresolved purchase. Record the next required approval and date. Keep verbal interest, acceptance of terms, and payment as separate observations.

Copy this pilot brief before recruiting

Complete one brief for each materially different workflow. Share the delivery commitments with the customer and keep internal cost assumptions alongside them.

Participant and recurring task:
Required inputs and exclusion criteria:
Current alternative and reason to change:
Customer outcome and deadline:
Baseline effort, quality, and measurement source:
Acceptance rule, reviewer, and required evidence:
Number of deliveries that must meet the rule:
Input format, volume limit, and deliverable:
Duration and work cycles:
Excluded work:
Manual involvement and customer responsibilities:
Approved data and who will review it:
Setup allowance:
Recurring delivery and support cap:
Pilot fee and payment terms:
Proposed continuation price and included service:
Buyer and required approvals:
Closing review date:
Evidence required to continue, revise, or stop:

For each delivery, keep a short log:

Account / cycle / scope version:
Output accepted? Reviewer and evidence:
Customer minutes: preparation / review / corrections:
Founder minutes: setup / delivery / support:
Manual changes to software output:
Direct costs:
Errors, extra requests, and scope changes:
Acceptance rule and effort cap met?

Your next task: make failure recognizable

Fill out the brief for one recurring customer task. Ask someone outside the project to read it and explain what would count as failure for both the customer outcome and your delivery effort. Tighten any rule they cannot apply.

Then bring the brief to one qualified prospect. Confirm that the outcome matters, the required inputs are available, and the proposed continuation terms are worth considering before you recruit more participants.