Frameworks for the work after the demo.

Operating systems

A compact library of questions and operating models for leaders who need AI adoption to become safer, clearer, and more useful work. Start with the work, test the next run, and make the human role real.

Relevance before technology

01

Find the specific work that loses context, repeats information movement, or sees warnings too late before deciding that AI is the answer.

Second-run proof

02

Judge a workflow by whether the next run needs less correction, has clearer exceptions, and produces a more usable outcome.

Oversight with authority

03

A reviewer needs a named role, visible evidence, enough capacity, and authority proportionate to the consequence of getting the work wrong.

Begin with the work, not an AI verdict.

01 / Relevance

In the second quarter of 2026, 40% of surveyed Canadian employer businesses said that AI was not relevant to the business. That result is real. It is not an employee-use census, and it cannot tell an operator whether AI is useful, safe, or repeatable in a specific piece of work.

The practical starting point is smaller than an entire company. Ask whether people regularly get information wrong or struggle to remember what happened before; whether they spend more than an hour a week moving information between places; and whether the business had the data and warning signs to avoid a problem but failed to connect them in time.

Those questions do not preselect AI. The honest answer might be better documentation, search, integration, ordinary automation, process repair, or no change at all. They locate the operating problem before a tool is asked to solve it.

Source context: Statistics Canada’s Q2 2026 analysis of business AI use and Jeremy Thomas’s July 2026 interview record.

Run it twice before you call it value.

02 / Second run

The first run of an AI workflow is a poor referendum on the idea. It is also a poor excuse for keeping a bad idea alive. A useful question is simpler: was the second run easier or better? Was the third?

Measure the whole loop: the time required to prepare the input, the meaningful work completed by the system, the correction and verification burden, the exceptions it created, the work quietly handed to another team, and the final outcome. A polished middle does not count if the cleanup queue simply moves downstream.

Then make one deliberate change: clarify the input, add evidence, define the acceptable output, narrow the workflow, or route a risky case to a person. Healthy implementation shows some learning—fewer corrections, less preparation, faster review, clearer ownership, or better handling of known exceptions. If the same defect returns after real changes, redesign the work or stop.

Source context: Jeremy Thomas’s July 2026 interview record; complementary-capability context from Statistics Canada’s analysis of AI adoption and productivity.

Give the human reviewer a job.

03 / Oversight

“A human reviews it” is not a control plan. Who is the person? What evidence can they see? Can they stop the action? What happens after they find a problem? Without answers, the human is a line in a slide deck, not meaningful oversight.

The minimum is modest: a named owner accountable for the consequence and visible evidence behind the output. Higher-consequence work may also need authority to pause or reject, an escalation route, enough time to investigate, and a way to feed known failures back into the workflow. The right level of control depends on consequence, uncertainty, reversibility, and observed failure—not on a universal checklist.

A useful reviewer does more than approve or reject. They explain why an edge case mattered, what signal should have revealed it, and whether the system should learn, narrow its scope, or hand that class of decision back to a person. The aim is not a larger approval log. It is fewer recurring defects.

Source context: Jeremy Thomas’s July 2026 interview record; NIST’s guidance on human-AI interaction and the Office of the Privacy Commissioner of Canada’s generative-AI principles.

Core question

If the tool works in a demo but fails in the business, what part of the operating system was missing?