ApproachServicesResearchCase StudiesAboutQ&ALet's Talk

Building an AI-native evidence system

A six-week engagement with a small online ESL school in China. The quality of the teaching was never the problem. The problem was that the quality was invisible to parents at the moment of the enrolment conversation.

124
evidence cards at handover, against a 100-card target
~300
cards four weeks after introduction: sustained adoption
3 / 3
contracted deliverables live, plus post-handover features
6 wks
from discovery to handover document

At a glance

Client
Go35 English, a small online ESL school operating in China (~150 students, 11 foreign teachers)
Engagement
Stage One pilot, one-month scope, three deliverables
Platform
Feishu Base with native AI Autofill; DeepSeek-V3.2 via Volcano Engine; Seedream 5.0 for image rendering
Result
Live system. 124 evidence cards submitted by handover against a 100-card target; approximately 300 cards sustained four weeks post-introduction. The system replaced the school's manual daily feedback forms.
Context
Post-regulation education market; cross-cultural build with an English-speaking consultant in a Chinese-first client environment

The problem

Go35 English is a small, quality-focused online school teaching English to Chinese young learners. Their foreign teachers are CELTA-qualified and their curriculum is drawn from serious educational publishers. Their commercial problem was not quality. It was that the quality was invisible to parents at the point of the enrolment or renewal conversation.

Referrals had weakened. Parents were comparing Go35 to cheaper platforms and asking why the price was higher. The sales team did not enjoy selling because they had no consistent way to answer that question. What happened in the classroom stayed in the classroom. Chinese teachers on the parent-facing side had to describe learning outcomes verbally, week after week, without structured evidence to hand.

The founder had tried to build a solution herself: a 39-page Feishu document, eleven tables, thirty-three planned automations. It was ambitious and it was stalled. What she needed was not more architecture. It was a working system her foreign teachers would actually use, and her Chinese teachers could actually turn into parent-facing evidence.

The engagement

Contracted as a one-month Stage One pilot, scoped deliberately narrow. Three components:

  • A positioning framework and objection-response scripts for the two most common parent objections: price, and time-to-visible-results.
  • A sales materials library populated with real classroom evidence from the month.
  • A foreign teacher collaboration workflow, targeting 100 evidence cards generated by teachers across the month.

One boundary was written explicitly into the contract: the consultant designs the system, the teachers generate the cards. This protected both sides. It forced the system to be genuinely usable, because teachers would abandon anything that wasn't, and it kept the consultant out of the position of writing content for the client.

Approach

Build for the person who has to use it

Four groups wanted four different things. The founder wanted a growth chart. The teachers wanted a form they could complete in under a minute on their phone between classes. The Chinese teachers wanted to see what had happened in a lesson without asking. Parents wanted a note that read like a person wrote it.

The failure mode of most SMB AI adoption is building for the person who signed the contract while quietly making the tool worse for the person who has to use it. So the teacher's experience came first, on a simple theory: if teachers used the system, everything else would follow. If they didn't, nothing else would matter. They did, so it did.

Ship early, then let real use write the specification

The first build was a working end-to-end system, deliberately without level calibration. As teachers used it, the specification improved through practice: the parent-note prompt was refined to match the Chinese-first register, teacher surveys reshaped the descriptors, and, following a proposal to management, the evidence cards replaced the manual daily feedback forms teachers had been writing, consolidating the workflow rather than adding to it. Only then did the second build add curriculum-calibrated level bands, informed by what real cards had revealed.

Transferable insight

Ship a working system early, learn what needs to change from real usage, then build the second version with the benefit of that learning. It saves weeks of arguing about specifications that were never grounded in real use.

Fit the client's context, not the capability leaderboard

The system was built inside the client's existing platform, on domestic AI models chosen for data-protection and integration reasons over the US models the team personally preferred. And two fragile edges were flagged in writing during the build, not after: a personal-account billing dependency, and an over-ambitious automated growth chart. Both later failed, exactly as flagged. Because they were named in advance, the response was recovery rather than crisis.

Timeline

May 2026

Discovery. The founder's self-built 39-page system: ambitious, stalled. Scoped down to a Stage One pilot focused on the school's most urgent commercial problem: enrolment.

Early June

Contract signed. Three deliverables scoped, with the consultant-designs / teachers-generate boundary written in explicitly.

Mid June

First build shipped in the client's own tenant. Teachers began submitting cards; the AI generated bilingual parent notes from day one.

Late June

Iteration through use: prompt register refined, descriptors revised from teacher surveys, workflow consolidated to replace manual daily feedback forms. Second build added CEFR-aligned level calibration. Growth chart rethought as a hybrid AI-plus-human model once real cards revealed per-card volatility.

Early July

System paused for several days when the AI provider's free token grant exhausted: the personal-account risk, flagged earlier in writing, materialising on schedule. Recovery via paid credit activation. Handover document drafted; card count 124 against the 100 target.

Mid July

Handover meeting. Three additional workflow features requested and delivered within the week, including a submission dashboard and automated notifications, plus a how-to guide for the parent-facing team.

Late July

Card image rendering designed and validated against real card data. Adoption metric at time of writing: approximately 300 evidence cards, four weeks after introduction.

Outcomes

The contracted deliverables

  • Narrative AI: positioning framework plus eight objection-response scripts against a contracted minimum of six, with new-parent and renewal tracks. Market research was conducted in Chinese-language sources so the scripts used native parent objection phrasing rather than translations of English templates.
  • Sales Materials Library: live, with the standout card pool exceeding 20 and four mini case studies built from teacher interviews, each reviewed by the interviewed teacher before publication.
  • Foreign Teacher Collaboration Workflow: live. 124 evidence cards at handover against a target of 100; approximately 300 four weeks after introduction, partly enabled by the workflow consolidation that replaced the manual daily feedback forms.

The system itself

  • Three linked tables plus a teacher submission form and role-specific views
  • Two AI fields per card, running on the platform's native AI Autofill
  • Level calibration through a 37-code curriculum mapping, editable by any admin without touching the AI logic
  • Bilingual output designed for parents to receive via WeChat
  • A hybrid AI-plus-human growth interpretation model

What this taught me about SMB AI adoption

Five patterns showed up here that show up almost everywhere. Each comes with the same practical takeaway I'd give any founder commissioning AI work.

1 · Personal-account dependency is the number-one silent failure mode

A production pipeline ran on one team member's personal account with a one-time free credit grant nobody knew was one-time. When it exhausted, everything stopped. Find these arrangements and fix them before build, not at 09:47 on a Monday when the tool stops working.

2 · Free credits get you to a pilot; production needs capital

Ask your consultant for a rough cost curve: what does month one look like on free credits, and month six at full volume? The free tier gets you a working pilot. Sustained operation needs a paid arrangement, and it's easier to plan for up front.

3 · Cross-cultural work needs a discipline single-language work does not

AI is genuinely useful for written artefacts and unavailable in live conversation. Budget conversation time as if the AI does not exist, and only count its contribution toward the written work.

4 · Outcome-based scoping is scope creep by design

Ask the consultant to itemise deliverables by build component — fields, prompts, views, documents — rather than by outcome. The consultant gets clearer scope, and you get a clearer picture of what the money buys.

5 · The artefact is not the goal; the conversation is

AI produces artefacts; artefacts feed conversations; conversations drive commercial outcomes. If your AI output will be used inside a human conversation, design it around what the person receiving it needs to feel or understand, and work backward.

Reflections

What I would take into every future engagement is the discipline of naming the fragile edges before they fail. The personal-account dependency. The expectation gap on the growth chart. Both were flagged in writing during the build. Both failed on schedule. Because they were named, the response could be recovery rather than crisis. That discipline, more than any specific technical skill, is what turned this engagement into a working system rather than a stalled one.

For any SMB founder considering an AI adoption engagement: the right question is not what the AI can do. The right question is who is going to use it, what conversations they need to have around it, and what the recovery looks like when a piece of it fails. If your consultant is comfortable talking about the failure modes before they happen, they are building for you.

Sound like the kind of build your business needs?

Every engagement starts with a discovery call: a straightforward conversation about your business and what you're trying to solve.