# Report Watch: three-pilot validation plan

**Planning template, not a payment request. Outcomes are unknown.** Public signup uses ChatGPT sign-in and an unpaid workspace preview, capped at 25 total retained accounts, including existing owner/pilot and held accounts. Public report monitoring requires confirmed live payment; the owner and up to three explicitly invited pilots retain separate admission. Live checkout and billing portal navigation are verified, but an actual purchase, mapped paid-customer entitlement and cancellation acceptance are not established by this guide. Check the current app status before proposing a paid evaluation. This document does not authorize outreach, spending or charges. The goal remains three paying pilots before broad investment, after the relevant paid-launch checks pass.

Use the [blank three-slot scorecard](/examples/pilot-scorecard.csv). Keep the original blank. Store completed copies privately, outside this website's `public` directory and source repository. The template contains no customers or results.

## 1. Choose a narrow fit

For each of the three slots, qualify one agency or small team with one existing daily or weekly CSV report, a named person who checks it, and a recurring deadline. Establish what a bad report costs them in actual review or correction effort; do not suggest a savings figure. Ask:

- What do you check now, how often, and what happened the last time a report was missing or unusable?
- Which column names, row range, date freshness, timezone, deadline and grace period would make a useful check?
- Who can authorize use of this report, and who will review the evidence during the pilot?

The fit is CSV receipt and configured-rule checking. It does **not** prove delivery to a client's inbox or storage, financial accuracy, or correctness of every business value. A missing-report incident means **no passing CSV was received for that deadline**, even if a failing file arrived. Background scheduling and email delivery have no promised SLA.

Keep within current limits: UTF-8 CSV, 1 MiB, 100 columns, 10,000 data rows, 100 UTF-16 units per header before and after normalization, and 16,384 UTF-16 units per data cell. A workspace allows 1,000 new checks per rolling 24 hours and 5,000 saved checks total. Failures count; identical retries consume no additional slot. The total cap does not clear with time; request export/retention review rather than retrying indefinitely. Owner/pilot and Solo allow five active reports; Agency allows 25. This experiment needs one report per pilot. Do not place separate customers together in the owner's workspace to bypass access or billing limits.

## 2. Obtain permission and agree on readiness

Before a participant starts, obtain explicit agreement on scope, monitoring access, data handling, evaluation dates, manual involvement, payment terms if applicable and withdrawal. Start with synthetic files, then sanitized reports with neutral column names; do not upload personal, confidential, regulated or unnecessary client data. Raw CSV rows are processed in memory, but column names, digests, check metadata, support messages and notification history are stored. Keep raw rows, secrets, credentials, contact details and billing identifiers out of the scorecard and support prompts.

A public page and a successful sign-in do not prove operational readiness. Record which parts of the agreed evaluation have passed these checks before starting its 14-day clock:

- Supported ChatGPT authentication, account enrollment/admission, monitoring entitlement, account isolation and dashboard upload work end to end. An unpaid public preview cannot run reports; an invited-pilot test does not prove the paid-account path. Test any external workflow separately before relying on it; manual uploads are not proof of unattended ingestion. Never share the owner's credentials.
- Scheduled missing-report detection, receipt validation and recovery are verified. If evaluating alerts, verify a **real queued branded email**, failure visibility and provider budget controls. Sender verification or a standalone test email alone is insufficient. If alerts remain unverified, agree on manual dashboard review and mark alert delivery unevaluated.
- The participant can obtain and inspect their workspace export, and withdrawal handling is understood. The existing browser check confirmed a download request, not the downloaded file contents. A deletion request places a hold; withdrawal keeps that hold until owner review. Neither means completed erasure or subscription cancellation. Review the available assisted-cleanup procedure and retention exceptions; do not promise immediate deletion or use sensitive customer data.
- Support expectations are explicit. Documentation answers and owner escalation are fallback paths. AI support is configured with a shared $5 monthly reservation allowance, at most 100 attempts across all workspaces; exhaustion falls back to documentation/owner review. New support conversations are capped at three per rolling day for unpaid accounts or 30 for paid/owner/invited-pilot operation. Verify the participant's escalation and owner-reply path, including any expected email; a successful owner-only test is not that proof.
- Before requesting payment: verify live checkout, signed event persistence/replay, entitlement activation only after confirmed payment, portal cancellation and the response to payment failure. Review the actual monthly offer and cancellation/refund-support terms with the participant. Sandbox acceptance, provider account activation, a trial, a zero-value invoice and an unmapped event receipt are not proof of paid customer access. Do not initiate a charge as part of this planning template.

Unverified parts stay outside the pilot's agreed scope. Record manual assistance and limitations; do not describe this beta as fully autonomous. Resolve retention/deletion and required provider checks before paid expansion. This guide does not itself satisfy any readiness check.

Owner verification on October 10, 2026: a support escalation and native queue flush sent one branded email from notifications@alerts.mayd-it.com, confirmed delivered by the email provider at 18:35:43 UTC. This proves that owner test only, not every pilot recipient, report alert or failure-notification path. The custom domain reportwatch.mayd-it.com has active domain and TLS status; provider configuration alone does not prove the current application release or customer ingestion.

Later October 10 verification: real Stripe-signed subscription-created, invoice.paid and subscription-deleted fixture events each returned204 and persisted three live receipts; duplicate cancellation preserved the same count and timestamps. The fixture was a canceled, unmapped zero-invoice trial. Live Solo checkout and the branded billing portal opened from the owner app and were exited without a purchase or change. Owner pilot access is separate from paid public entitlement. These transport/navigation results do not satisfy the mapped paid-customer gate above. Report alerts are also bounded: each deadline can queue one quality-failure and one recovery message, with missing-report and support messages separate. Later revisions still update incident/check evidence, but do not repeat those quality/recovery emails. Maintenance attempts at most five emails per run; queueing and a successful sweep do not prove inbox delivery.

## 3. Measure a baseline, then a 14-day pilot

Before enabling the pilot, record manual review minutes across at least three comparable scheduled report periods. Prefer a timed prospective baseline. If using recalled or historical estimates, label them and do not treat them as measured savings. Include reviewing the report and investigating issues; exclude unrelated report production. Record the same work during the pilot, and track Mayd-it operator/support minutes separately.

Agree on a 14-day evaluation and run four controlled cases with synthetic data in a separate test monitor: a valid fresh file passes; a required-column failure is flagged; stale dates are flagged; and a deadline with no passing submission produces the expected incident after grace and a scheduled sweep. Verify expected evidence, not merely an HTTP success. Do not inject failures into a customer's live workflow.

During the evaluation, manually compare every due report period with the source workflow and saved evidence. Count distinct due periods, not upload attempts; exclude paused periods and deadlines whose grace has not elapsed. Record false alerts and known problems the configured checks missed. A passing recovery received late does not turn that period into an on-time passing period. Keep dated observation notes in a separate private record, referenced without customer identifiers.

At day 14, ask the participant to explain one result, show how they used it without prompting, and say whether they want to continue. Weekly reports usually yield only two cycles in 14 days: that is insufficient evidence for the proposed four-cycle gate below. Mark the result inconclusive and agree on a longer evaluation before extending it. Do not manufacture extra live cycles.

## 4. Make a decision from evidence

These are **proposed thresholds to agree before testing**, not performance claims:

- Technical: all four controlled cases behave as expected; zero observed false alerts and zero known missed in-scope problems across at least 10 daily or four weekly due periods. A small clean sample is not proof of future reliability.
- Usefulness: the participant can interpret the evidence, returns without a reminder, and requests continued use. Compare review minutes per due period with a comparable timed baseline; a 25% reduction is a test hypothesis. Include operator effort when judging whether the service can become sustainable. Estimated baselines or too few cycles make the efficiency result inconclusive.
- Demand: discuss **$49/month for up to five reports** or **$149/month for up to 25 reports** as price hypotheses, only when relevant to the participant. Ask which, if either, they would choose and why. Interest, a verbal commitment and a sandbox checkout are not payment. Count a paying pilot only after an authorized live payment is confirmed in the payment provider; record a private evidence reference, never payment credentials.

For each slot, record `continue`, `revise`, `stop` or `inconclusive` and why. Fix correctness or access failures before expansion. Three interested teams do not meet the three-paying-pilot goal; no conclusion about revenue or autonomous operation is available until real evidence supports it.

## Scorecard field definitions

Leave unknown values blank; use `0` only for a measured zero. Dates use UTC ISO 8601. Numeric fields contain plain numbers, not currency symbols or formulas. The CSV is a manual template, not an automated score or customer database.

| Fields | What to record |
| --- | --- |
| `pilot_id` | Fixed neutral slot identifier: pilot-01, pilot-02 or pilot-03. |
| `qualification`, `consent_recorded`, `readiness_confirmed` | Qualification: fit / not_fit / uncertain. Consent and launch gates: yes / no; leave unknown blank. Keep underlying permission records privately. |
| `cadence`, `baseline_method` | daily / weekly; timed / historical_estimate / recalled_estimate. |
| `baseline_periods`, `baseline_review_minutes` | Comparable baseline due-period count and total manual review minutes. |
| `pilot_start_utc`, `pilot_end_utc` | Actual evaluation dates, populated only when it starts/ends. |
| `due_periods`, `passing_periods` | Evaluated due periods; those with at least one passing CSV received by that period's deadline plus grace. Count each period once. |
| `controlled_cases_passed` | Number of the four synthetic cases with verified expected behavior, 0–4. |
| `false_alerts`, `missed_known_issues` | Manually confirmed false incidents and missed in-scope issues. Keep the comparison evidence privately. |
| `review_minutes`, `operator_minutes` | Total participant review/investigation time and separate Mayd-it assistance time during evaluation. |
| `evidence_understood`, `return_use`, `continuation_interest` | yes / no based on the final conversation and observed use; do not infer from a page view. |
| `monthly_price_tested_usd`, `price_response` | 49 or 149 if discussed; acceptable / too_high / not_needed / undecided. This records an interview, not a charge. |
| `payment_status`, `payment_evidence_ref` | not_requested / interest_only / committed_unpaid / sandbox_only / paid / refunded. For an actual transaction, use a reference to a private verification record; no provider/customer IDs or receipt URLs here. |
| `decision`, `decision_note` | continue / revise / stop / inconclusive plus a short factual reason without customer data. |
