OPERATIONS / SYSTEM DESIGN
Designing an Operating System for Scale
The business did not simply need more SOPs. It needed a clearer answer to three questions: what should the team remember, what should they decide, and what should the workflow catch for them?
Monthly GMV during the scaling period
Approximate increase in order volume
Problematic orders per month
THE SCALING PROBLEM
The business scaled faster than the way it carried knowledge.
Informal strengths became structural dependencies.
At roughly $150K in monthly GMV, informal coordination was an advantage. Experienced employees could remember exceptions, teach by example, and resolve unusual cases directly.
Above $1M, the same shortcuts became dependencies. New people could copy visible actions without understanding the decisions behind them; an exception could cross more hands than the message explaining it; management became the fallback for both routine questions and genuinely risky judgment calls.
THE DIAGNOSIS
The question was not “what can become an SOP?” It was “where should each kind of complexity live?”
More documentation would not solve every failure. Stable knowledge could be codified; repeatable decisions could be delegated; and preventable errors could be blocked by the workflow. Contextual judgment still had to remain with people.
Knowledge
Move the stable core of strong execution out of shadowing and into a system the next operator can use.
Decisions
Push familiar, low-risk calls to the frontline; reserve management attention for novelty, ambiguity, and downside.
Errors
Put information and evidence inside the physical flow so prevention does not depend on memory or uninterrupted attention.
THE RESPONSE
One diagnosis. Three different interventions.
I did not apply one universal playbook. Each workflow failed for a different reason, so each required a different balance between standardization and human judgment.
Tacit knowledge → Repeatable execution
Livestream Operations
The failure was not a lack of talent. Strong hosts and directors had learned to make dozens of small decisions intuitively, but the business had no reliable way to transfer that logic to a larger team.
Ambiguous judgment → Structured decision rights
Customer Service
The bottleneck was not response speed alone. Routine cases reached management too often, while a less experienced employee could still make the wrong call on a financially or reputationally sensitive dispute.
Informal coordination → Embedded controls
Fulfillment
At low volume, a verbal reminder could follow an order. At roughly ten times the volume, the order moved across more people than the reminder did—and accuracy depended on someone catching the mismatch at the end.
01 / LIVESTREAM OPERATIONS
Codifying the floor—not scripting the ceiling.
I standardized what every session had to get right: presentation sequence, product roles, director responsibilities, and feedback. I deliberately left pacing, energy, storytelling, and live adaptation to the operator.
BEFORE
Learn by watching.
AFTER
Learn through structure.
Minimum viable consistency
A core presentation flow created a reliable baseline. Situational modules helped newer hosts recognize contexts such as scarcity, value, or low room energy without prescribing a word-for-word performance.
Feedback as memory
Host profiles and post-stream reviews preserved specific coaching across sessions, so improvement no longer depended on whether the same manager happened to be present.
Checks at the point of work
A visible self-check system surfaced the few responsibilities most likely to disappear during a busy stream. It replaced retrospective reminders with prompts during execution.
Strategy made physical
Products were classified as traffic drivers, proven sellers, or scarce inventory and separated into physical zones. The room itself began carrying part of the product-selection logic.
02 / CUSTOMER SERVICE
Moving routine decisions outward—and risk inward.
The objective was not to eliminate discretion. It was to stop treating every customer issue as equally uncertain: documented cases could move quickly, while financial and reputational risk stayed with experienced decision-makers.
Frontline
Pre-sale and standard after-sale cases.
Price · Condition · Shipping · Accessories · Basic order issuesExperienced Staff
Risk cases within documented decision boundaries.
Condition disputes · Return requests · Existing risk scenariosManagement
Novel, ambiguous, or high-risk exceptions.
New scenarios · Policy decisions · Unclear resolution boundariesTHE LEARNING LOOP
A new exception may consume management judgment once.
A recurring pattern should not consume it forever.
03 / FULFILLMENT
Moving control into the physical flow.
As order volume increased roughly tenfold, “be more careful” was not a control. Critical information had to travel with the order, and evidence had to be available where each handoff occurred.
INFORMATION TRACEABILITY
Make exceptions visible.
Special requests moved from verbal communication to written records. Simple product-tag notation kept the exception attached to the item instead of depending on the original employee being available to explain it again.
PREVENTIVE CONTROL
Make errors harder.
A quick reference system surfaced accessory images before shipment, while SKU photos and markings on both boxes and shipping labels created independent evidence at the handoffs where a mismatch could enter the process.
MEASURED RESULT
Problematic fulfillment orders per month.
The reduction occurred while order volume was roughly ten times higher. This was the clearest measurable result of the redesign; the livestream and customer-service systems addressed different constraints, so I did not force them into the same KPI story.
OPERATING PRINCIPLES
The choices behind the systems.
These are not universal rules. They are the tests I used to decide what to codify, what to delegate, and what to leave to judgment.
Standardize the floor, not the ceiling
Codify what acceptable execution always requires, then leave room for strong operators to outperform it.
Put information at the point of action
A reminder in someone's memory is weaker than a visible cue inside the workflow where the decision happens.
Escalate by risk, not hierarchy
Known, low-risk cases should move without management. Novelty, ambiguity, and downside should trigger escalation.
Make control earn its friction
Every checkpoint adds work. I added one only when the cost or frequency of failure justified slowing the process down.
REFLECTION
The operating system should absorb repetition—not judgment.
Scaling meant deciding which complexity the organization could carry itself and which still deserved human attention. The result was not a larger stack of SOPs; it was a business that relied less on memory, escalated more deliberately, and made preventable errors harder to create.