← Back to projects

Designing an Operating System for Scale

The business did not simply need more SOPs. It needed a clearer answer to three questions: what should the team remember, what should they decide, and what should the workflow catch for them?

BUSINESS CONTEXT$150K → $1M+

Monthly GMV during the scaling period

ORDER VOLUME~10×

Approximate increase in order volume

FULFILLMENT ISSUES5–10 → ~1

Problematic orders per month

The business scaled faster than the way it carried knowledge.
Informal strengths became structural dependencies.

At roughly $150K in monthly GMV, informal coordination was an advantage. Experienced employees could remember exceptions, teach by example, and resolve unusual cases directly.

Above $1M, the same shortcuts became dependencies. New people could copy visible actions without understanding the decisions behind them; an exception could cross more hands than the message explaining it; management became the fallback for both routine questions and genuinely risky judgment calls.

More volume→More people→More handoffs→More complexity

The question was not “what can become an SOP?” It was “where should each kind of complexity live?”

More documentation would not solve every failure. Stable knowledge could be codified; repeatable decisions could be delegated; and preventable errors could be blocked by the workflow. Contextual judgment still had to remain with people.

REMEMBER

Knowledge

Move the stable core of strong execution out of shadowing and into a system the next operator can use.

JUDGE

Decisions

Push familiar, low-risk calls to the frontline; reserve management attention for novelty, ambiguity, and downside.

CATCH

Errors

Put information and evidence inside the physical flow so prevention does not depend on memory or uninterrupted attention.

One diagnosis. Three different interventions.

I did not apply one universal playbook. Each workflow failed for a different reason, so each required a different balance between standardization and human judgment.

01

Livestream Operations

The failure was not a lack of talent. Strong hosts and directors had learned to make dozens of small decisions intuitively, but the business had no reliable way to transfer that logic to a larger team.

STANDARDIZED — presentation floor, product roles, director checksKEPT HUMAN — pacing, energy, storytelling, situational adaptation
02

Customer Service

The bottleneck was not response speed alone. Routine cases reached management too often, while a less experienced employee could still make the wrong call on a financially or reputationally sensitive dispute.

STANDARDIZED — known cases, resolution ranges, escalation triggersKEPT HUMAN — novel, ambiguous, and high-risk exceptions
03

Fulfillment

At low volume, a verbal reminder could follow an order. At roughly ten times the volume, the order moved across more people than the reminder did—and accuracy depended on someone catching the mismatch at the end.

STANDARDIZED — exception records, visual evidence, SKU checkpointsKEPT HUMAN — final verification when evidence conflicts

Codifying the floor—not scripting the ceiling.

I standardized what every session had to get right: presentation sequence, product roles, director responsibilities, and feedback. I deliberately left pacing, energy, storytelling, and live adaptation to the operator.

Learn by watching.

Individual experience↓Intuitive decisions↓Shadowing & imitation↓Inconsistent execution

Learn through structure.

Core standards+Situational playbooks+Inventory logic+Feedback loops

Minimum viable consistency

A core presentation flow created a reliable baseline. Situational modules helped newer hosts recognize contexts such as scarcity, value, or low room energy without prescribing a word-for-word performance.

Feedback as memory

Host profiles and post-stream reviews preserved specific coaching across sessions, so improvement no longer depended on whether the same manager happened to be present.

Checks at the point of work

A visible self-check system surfaced the few responsibilities most likely to disappear during a busy stream. It replaced retrospective reminders with prompts during execution.

Strategy made physical

Products were classified as traffic drivers, proven sellers, or scarce inventory and separated into physical zones. The room itself began carrying part of the product-selection logic.

Moving routine decisions outward—and risk inward.

The objective was not to eliminate discretion. It was to stop treating every customer issue as equally uncertain: documented cases could move quickly, while financial and reputational risk stayed with experienced decision-makers.

LEVEL 01

Frontline

Pre-sale and standard after-sale cases.

Price · Condition · Shipping · Accessories · Basic order issues
→
LEVEL 02

Experienced Staff

Risk cases within documented decision boundaries.

Condition disputes · Return requests · Existing risk scenarios
→
LEVEL 03

Management

Novel, ambiguous, or high-risk exceptions.

New scenarios · Policy decisions · Unclear resolution boundaries

A new exception may consume management judgment once.
A recurring pattern should not consume it forever.

New case→Escalate→Resolve→Define principle→Update SOP

Moving control into the physical flow.

As order volume increased roughly tenfold, “be more careful” was not a control. Critical information had to travel with the order, and evidence had to be available where each handoff occurred.

Make exceptions visible.

Special requests moved from verbal communication to written records. Simple product-tag notation kept the exception attached to the item instead of depending on the original employee being available to explain it again.

Make errors harder.

A quick reference system surfaced accessory images before shipment, while SKU photos and markings on both boxes and shipping labels created independent evidence at the handoffs where a mismatch could enter the process.

Product SKU→SKU photo→Box marking→Shipping label→Final check
5–10→~1

Problematic fulfillment orders per month.

The reduction occurred while order volume was roughly ten times higher. This was the clearest measurable result of the redesign; the livestream and customer-service systems addressed different constraints, so I did not force them into the same KPI story.

The choices behind the systems.

These are not universal rules. They are the tests I used to decide what to codify, what to delegate, and what to leave to judgment.

01

Standardize the floor, not the ceiling

Codify what acceptable execution always requires, then leave room for strong operators to outperform it.

02

Put information at the point of action

A reminder in someone's memory is weaker than a visible cue inside the workflow where the decision happens.

03

Escalate by risk, not hierarchy

Known, low-risk cases should move without management. Novelty, ambiguity, and downside should trigger escalation.

04

Make control earn its friction

Every checkpoint adds work. I added one only when the cost or frequency of failure justified slowing the process down.

The operating system should absorb repetition—not judgment.

Scaling meant deciding which complexity the organization could carry itself and which still deserved human attention. The result was not a larger stack of SOPs; it was a business that relied less on memory, escalated more deliberately, and made preventable errors harder to create.