Heard us on
The AI Daily Brief Podcast?

Heard us on The AI Daily Brief Podcast? For AI, we're all in on AWS. Let's build AI teammates for your enterprise.

Part 1: The AI Productivity Paradox – Your Teams are Already Changing Their Jobs, Have You Noticed? 

AI has sped up work, but few companies can point to organization-wide return on investment (ROI). Everyone is calling this a paradox. This three-part series names the three things buying the tools never fixes on its own, the people, the organization, and the foundation, and what to do about each. Read all three to find where your own gap is hiding.

Part 1: The People | Part 2: The Organization | Part 3: The Foundation

Somewhere in your company right now, a handful of people have become several times more productive with AI. Not because of the training program everyone had to take or the Copilot licenses that were purchased. On their own, on real work, because it made their week better.  

If you asked them, they would tell you it feels like work that used to take them days, now completes in minutes. They are in every organization. 

But your company’s productivity numbers didn’t move. That gap is the paradox in its most personal form.  

Atlassian’s 2026 State of Teams research puts a number on it: 89% of executives say AI has increased the speed of work, but only 6% can point to organization-wide ROI. Individual speed is everywhere. Organizational results are rare. The interesting question isn’t whether AI works; your people have already answered that. It’s why individual 10x doesn’t add up to organizational 10x. 

Where the gains go 

The gains are real. They’re just trapped. Three things trap them, and none of them is technical. 

First, the gains live in individuals. The prompts, the workflows, the personal knowledge base, the judgment about when to trust the AI output and when not to, all of it sits in someone’s chat history and someone’s head. When that person is on vacation, the gain is on vacation. When they leave, it leaves with them. 

Second, nobody can see the gains from the outside. The report still lands on Friday. The analysis still arrives before the meeting. The work looks identical to what it looked like a year ago. People are using the time they saved to do more of the old job, or to go home at a reasonable hour. Neither shows up in a productivity dashboard. 

Third, people have a reason to keep quiet. The honest reaction to suddenly being several times faster is not pride. It’s something closer to this feels like we shouldn’t be doing this. If the work that used to take a week now takes an hour, what does that say about the week? About the role? About the headcount? Your most proficient people may be under-reporting, not over-claiming. The paradox has a visibility problem before it has anything else. 

What we learned running it on ourselves 

This summer we ran an experiment at Robots & Pencils called Robocon. Four weeks, company-wide, across seven cross-functional pods we were given challenges to build and publish working AI skills to a shared internal library, then show them off at a live event. No course. No training. Not just engineers, but everyone in the company. Build something real and put it in front of people. 

Two lessons came out of it, and neither was the one we expected. 

The first was that the barrier wasn’t skill. Teams lost hours, in some cases days, getting environments, repositories and file-share access to work before they could build anything. Once these hurdles were cleared, people began creating at an alarming rate. Everyone in the organization from marketing, sales, product, project managers, design, engineering, came out the other side with real AI proficiency. The expensive part of proficiency wasn’t teaching people. It was clearing the path so they could do the work. 

The second lesson landed harder. The library filled up fast, and plenty of good work shipped and then sat there, because nobody’s job was to turn what one person built into something the rest of the company used. Not everything built was worth keeping, and the pieces that were, had to be deliberately identified, hardened and surfaced to the rest of the company. Sorting them was a separate act, done by different people, with a different skill. 

That’s the pattern, and we think it generalizes. The event produced proficiency in individuals. Amplifying it across the organization was a second job, and it didn’t happen on its own. 

Roles aren’t redesigned. They are discovered.  

Most advice about AI and the workforce runs top-down: redesign the roles, then deploy the tools. In our experience it works the other way around. People pick up the tools, use them to create value for themselves first, and the role starts to change underneath them. The job description is the last thing to move. AI is moving from individual productivity to multiplayer, collaborative productivity.  

That’s the good news, because it means the change is already happening in your company. The challenge is that discovery doesn’t distribute itself. A role that has quietly changed in one person’s hands stays there unless someone does three things on purpose. 

Find it. Who has already changed how they work? You won’t learn this from a survey; you learn it by asking a different question: not “are you using AI?” but “what part of your job have you automated?” People who have stopped doing something manually are the ones whose role has evolved. 

Harvest it. Turn what one person does into something others can use. Not a training deck or show and tell, but the actual workflow, the actual prompts, data and tools they have connected, and the actual judgment calls about when the output can be trusted. Then be honest about what holds up. Our own library taught us that a lot of good individual work is not, in fact, reusable. The discipline is in keeping what is. 

Make it official.  Redefine the role around what changed and retire the old work. Skip this and people end up twice as fast in a job still defined as if they weren’t. The time saved gets refilled with more of the old job, and the productivity number never moves. 

We did this to ourselves. A client team that used to be seven to ten people is now three to five. AI didn’t take the seats — the people kept them and got more done. Each person now works with AI the way they’d work with a strong analyst: it drafts, researches, checks and runs the routine parts, and they spend their time on the judgment calls. Once that was true in practice, we rewrote how we define what a team is on paper: who’s on it, what each role owns, what nobody does by hand anymore. 

The job nobody has hired for 

Boris Cherny, who leads the Claude Code team at Anthropic, has described how engineering and product roles on his team have melted into six archetypes: the Prototyper who generates ideas most of which don’t ship; the Builder who turns a validated idea into production; the Sweeper who simplifies and removes; the Grower who iterates for scale; the Maintainer who owns the mature system; and the Orchestrator, who knows which of the other five a team should be in, and when to move between them. He notes the Orchestrator is rarer than the other five combined, and that it has no established job title. Most people are doing 2 or 3 of these archetypes today. 

We think that last role is the one the productivity paradox is missing. The Orchestrator is the person who does the finding, harvesting and socializing to teams. They notice that a role has changed before the org chart does.  

Nobody at this table has hired one. Most organizations don’t know it’s a job. But somebody in your company is already doing a version of it informally — the person other people go to when they want to know how so-and-so got that done so quickly. That person is your first Orchestrator. They just don’t have the title yet. 

Questions to take back 

Your people are already changing their jobs. The organizations that get results from AI aren’t the ones that designed the change from the top. They’re the ones that noticed it from the bottom and made it official before it evaporated. 

Two questions for your leadership team this week: 

Up next in this series, Part 2: Organization

Individual proficiency is the easy part. The harder part is turning it into something the organization can run on without you in the room. In Part 2, we break down the three structural mistakes that stall AI pilots before they reach production, and what to build instead. Read Part 2: Common mistakes in AI org design (and how to fix them).

Sources referenced: Atlassian, State of Teams 2026 (89% / 6% figures). Boris Cherny, public posts on role archetypes (X.com). 


About the Author

Brendan Flynn is SVP, Strategist at Robots & Pencils where he heads industry strategy within the Generative & Agentic AI Studio.


Key Takeaways

FAQs

What is the AI productivity paradox?

It is the gap between how fast people say AI has made their work and how rarely that speed shows up in company-wide results. Atlassian’s 2026 State of Teams research found 89% of executives report faster work from AI, while only 6% can point to organization-wide ROI.

Why don’t individual AI productivity gains show up in company-wide numbers?

The gains stay trapped in three places. They live in one person’s prompts and judgment instead of a shared process, they are invisible to standard reporting because the output looks the same as before, and people often stay quiet about how much faster they have become.

What did Robots & Pencils learn from running Robocon, its internal AI build event?

The barrier to AI proficiency was never skill. Once people had access to the environments and tools they needed, proficiency spread quickly across the company. The harder problem was turning one person’s work into something the rest of the organization could use.

What three steps turn individual AI proficiency into organizational capability?

Find who has already changed how they work, harvest what they are doing into a reusable workflow, and make the change official by redefining the role and retiring the old work it replaced.

What is an AI Orchestrator?

A term from Anthropic’s Boris Cherny for the person who notices when a role has changed because of AI and helps spread that change across a team. Most organizations already have someone doing this informally, without the title.

Part 2: The AI Productivity Paradox – Common Mistakes in AI Organization Design (and how to fix them) 

AI has sped up work, but few companies can point to organization-wide return on investment (ROI). Everyone is calling this a paradox. This three-part series names the three things buying the tools never fixes on its own, the people, the organization, and the foundation, and what to do about each. Read all three to find where your own gap is hiding.

Part 1: The People | Part 2: The Organization | Part 3: The Foundation

You have a few AI pilots that have worked. The results are promising, and the business case makes sense. You spin up a dedicated AI team, give them a mandate, and hand off the work. 

Six months later, nothing has shipped to production. The business side says the AI team doesn’t understand the constraints. The AI team says the business doesn’t understand what’s possible. Meanwhile, the productivity gains that looked obvious in the pilot have vanished. 

It’s the most common failure pattern we see, and it’s rarely a technical one. It’s organizational. 

After working through this with clients across finance, manufacturing, and operations, we’ve landed on three structural mistakes that kill AI initiatives before they scale. Here’s what they look like, and what we’ve found works. 

Mistake 1: Centralizing AI Decisions Away from the Business 

The error: You create an AI Center of Excellence or a dedicated AI team, and suddenly every AI decision routes through them. They’re the gatekeepers. The business waits. 

Why it fails: The people closest to the problem, the ops manager, the finance lead, the customer service director, can’t move fast. They hand requirements to the AI team, the AI team interprets them, requirements get misunderstood, timelines slip, and by the time something ships the business context has already shifted. 

What works instead: Push AI decisions closer to the business. You still need shared standards — how AI workflows get evaluated, how decisions get logged, how risk is governed — but the business unit lead, who owns the outcome, should decide when an AI workflow is ready to go live in their function. 

We made the same call inside our own delivery organization. Instead of standing up one central AI practice that every client team has to route through, we organized around small cross-functional pods, each aligned to a specific client, each deciding for itself how AI gets used in that engagement. A pod answers to shared standards, a hiring bar, a set of proficiencies, not to a gatekeeper reviewing its every move.  

Mistake 2: Treating Upskilling as Training, Not Proficiency Building 

The error: You run a one-week training program. Everyone learns to use Claude, Copilot, or ChatGPT. Then you expect productivity to jump. 

Why it fails: Training teaches how to use a tool. It doesn’t build organizational proficiency. The operations manager learns Claude syntax, goes back to their desk, and does the same job exactly the way they used to, just a little faster. Productivity gains are marginal, executives see no ROI, and everyone quietly concludes AI didn’t work. 

What works instead: Build proficiency through doing, not classrooms. 

Start by asking what this person stops doing, what new responsibilities they take on, and how their relationship to the work changes. Answer those honestly and you’ve effectively restructured how that function operates, at which point upskilling stops being a training event and becomes proficiency built through doing the work. 

A client we worked with had an immediate unlock the first time we had them use an orchestrated workflow we built for their forecasting process. The response was emphatic, “This will replace 90% of the meetings we have, I can simply ask AI questions about my forecast, and it knows my entire book.”  

Zero training involved, they got it immediately. That is powerful. This frees them up people to do what they do best, build relationships to close deals.  

Mistake 3: No Feedback Loop Between Operations and AI 

The error: The AI team builds an AI workflow, the business puts it to work, and six months later nobody can say whether it’s moving the outcomes that matter. 

Why it fails: Without real feedback you can’t improve: performance drifts, edge cases pile up, the business loses confidence, and the AI team never learns what’s needed. 

What works instead: Build operational feedback into the rollout from day one. Who’s watching how it’s performing against the outcomes you care about? Who surfaces problems? How fast can you respond? 

On another engagement, we skipped the weekly-sync approach entirely and went tighter. The day two account managers first tried the assistant live, on their own real data, their feedback went straight into a ranked list for engineering before the day was over. Four fixes made the build before the end of the same day the feedback was received. A handful of other requests got logged as legitimate but not urgent. A couple of ideas got an explicit “not this round,” with the reasoning written down so nobody had to relitigate it later.  

The Pattern Underneath 

These aren’t technology problems, they’re proficiency problems: companies think they’re buying an AI system when what they’re building is organizational capability. 

The fix comes down to three structural shifts, and they reinforce each other.  

Push AI decisions closer to the business, because feedback only flows fast when the people doing the work own the decision to change it.  

Redefine roles and workflows through doing instead of training, because software that keeps evolving forces people to build proficiency in real time, and that’s where the actual transformation happens.  

Close the feedback loop, then keep it tight, because iteration speed is the discipline: every cycle, the system improves, and the organization learns alongside it. 

When those three things hold, AI proficiency becomes the muscle of how the organization operates day to day, and the productivity gains stick instead of fading out after the pilot. None of it requires a heroic AI team. It requires distributed decision-making with clear guardrails around it. 

If your AI tools are live but proficiency, and the productivity gains that are supposed to come with it, still haven’t shown up, this is the first place to look. 

Questions to take back 

Up next in this series, Part 3: Foundations

Fixing the organization gets you further, but it does not answer the harder question underneath it. Once the model itself becomes a commodity, what should you actually own? In Part 3, we lay out the five-part foundation every company needs, whether it builds its own AI or simply buys it. Read Part 3: Own the foundation, rent the model.


About the Author

Brendan Flynn is SVP, Strategist at Robots & Pencils where he heads industry strategy within the Generative & Agentic AI Studio.


Key Takeaways

FAQs

Why do AI pilots that work so well often fail to scale?

Because the failure is usually organizational, not technical. If AI decisions route through one central team, if training substitutes for real proficiency, or if there is no feedback loop back to the business, the gains a pilot proved rarely carry into production.

Should a company centralize its AI decisions in one team?

No. Centralizing every AI decision in one gatekeeping team creates a queue the business waits behind. Shared standards should be centralized. The decision to put a specific AI workflow into production should sit with the business unit that owns the outcome.

What is the difference between AI training and AI proficiency?

Training teaches someone to use a tool. Proficiency changes how they do the job. A short training session on a chatbot rarely moves productivity. Asking what a role stops doing and what it takes on instead, then building that into daily work, does.

What happens without a feedback loop between the business and the AI team?

Performance drifts, edge cases pile up, and nobody can prove months later whether the AI workflow is helping. A tight feedback loop, where real users test the workflow on real data and issues go straight into a ranked list for the team building it, catches problems while they are still cheap to fix.

What three changes help AI initiatives scale past the pilot stage?

Push AI decisions closer to the business unit that owns the outcome, build proficiency through real work instead of classroom training, and close the feedback loop between operations and the team building the AI, then keep it tight.

Part 3: The AI Productivity Paradox – Own the Foundation, Rent the Model 

AI has sped up work, but few companies can point to organization-wide return on investment (ROI). Everyone is calling this a paradox. This three-part series names the three things buying the tools never fixes on its own, the people, the organization, and the foundation, and what to do about each. Read all three to find where your own gap is hiding.

Part 1: The People | Part 2: The Organization | Part 3: The Foundation

Count the AI tools running in your company right now. The ones actually in use. The Copilot licenses, Enterprise Claude accounts, Salesforce Agentforce. The pilot the finance team built with a contractor. The thing the marketing group pays for on a corporate card. The internal assistant somebody in operations put together over a long weekend. 

If you got past ten, you’re typical. If you can name who owns each one, what it costs, and whether it’s working, you’re rare. Atlassian’s 2026 research found that only 6% of executives can point to organization-wide AI ROI, and the reason isn’t that the tools don’t work. It’s that there is nothing to measure them with. Twelve tools, twelve logins, twelve vendors, no shared memory. Every pilot starts from zero. The company gets faster at individual tasks and no smarter as an organization. 

Why this matters more than it did a year ago 

Something shifted in the middle of 2026 that most executive teams haven’t fully absorbed: the frontier models became close to interchangeable. The gap between the leading providers narrowed to the point where, for most business tasks, which model you use matters far less than what you’ve built around it. The models are becoming commoditized. Prices fell. Switching got easier. The model has become something you rent. 

That changes what the durable asset is. When the model is a commodity, the thing you wrap around it, your business context, the guardrails, your way of knowing whether it’s working, is the only part that appreciates. 

The question for executive teams isn’t “which AI should we buy?” It’s “what should we own, regardless of what we buy?” Engineers call this the harness. We’ll call it the foundation. You rent the model. You own the foundation. 

What the foundation is, in plain terms 

Strip the architecture diagrams away and the foundation is five things. None of them is exotic. Most companies have a version of each for their financial systems and none for their AI. 

A shared definition of your business. What “a customer” means in your company. What “an order” is, what “a batch” is, what “on time” means, and the history of each, because most systems of record overwrite yesterday and remember nothing. Written down once, in a form every tool can use, so the finance team’s assistant and the operations team’s forecast are talking about the same thing. Without this, every AI tool learns your business from scratch, and each one learns it differently. 

One door. Every AI request in the company goes through a single point, so cost, usage and risk show up in one place. Without this, you cannot answer the CFO’s question about what AI is costing you, and you cannot swap vendors without rewriting everything that touches them. 

A register. A list of every AI tool running in your company, bought or built, each with a named owner. Without this, you don’t know what you’re governing. Most companies discover the size of their AI footprint the first time something goes wrong. 

A way to know if it’s working. Before anything goes live, and every week after. Not a vendor’s demo; your own test, on your own data, against what happens. Without this, “the pilot worked” is an opinion, and six months later nobody can say whether the thing in production still does. This is your evaluation framework. 

Approval in proportion to risk. Who signs off on what, at what level of consequence, with the same rules for the tools you bought as for the tools you built. Without this, governance either doesn’t exist or exists as a committee everything waits behind. 

That’s it. Five things. If you have them, you can add the eleventh tool in a week and know what it costs and whether it works by the second week. If you don’t, the eleventh tool is another island. 

A nervous system, not a headquarters 

Here is where executives get worried, and rightly. “Own the foundation” sounds like “centralize AI,” and centralizing AI is one of the fastest ways to fail. We wrote about that in the second article in this series: the AI Center of Excellence that becomes the place where the business goes to wait. 

The foundation is not that. Think of it as a nervous system, not a headquarters; it’s the core of what makes your business, yours. A nervous system doesn’t decide where the hand goes. It makes sure the hand can feel, that the signal gets back to the brain, and that the whole body knows what the hand just learned. The center owns the foundation, the shared definitions, the one door, the register, the tests, the rules. The business units own the decisions: what to build, what to buy, whether it goes live in their operation. Standards are shared. Decisions are local. 

Get this distinction wrong in either direction, and your AI initiatives won’t scale. Centralize all decision making and you have a bottleneck. Decentralize your foundation and you have twelve islands with twelve definitions of a customer. The companies getting results have done the unglamorous thing: they’ve been disciplined about what’s shared and explicit about what’s not. 

The tools you already bought count 

Most organizations are not in the business of building an agent-building program. You’ve bought licenses. That doesn’t exempt you from defining the foundation; it’s the strongest argument for it. 

Governance that only covers what you build misses most of what you run. The Copilot seats, the vendor’s embedded AI, the SaaS tool that quietly added a model last quarter: all of it consumes your data, produces outputs your people act on, and costs your organization money, and almost none of it shows up on the same budget line item as the internal pilot.  

The register is how it gets there. The one door is how you see what it costs, the tokenomics. The test is your evaluation framework to find out whether the expensive bundle you renewed in January is moving you closer to your north star. 

This is the accountability answer for a company that has bought AI and can’t see the return. You don’t need to have built anything to need a foundation. You need one because you bought things. 

What it looks like when it’s done right 

One company we worked with, a manufacturer with operations across several regions, had systems of record that kept no history. Each month overwrote the last. That’s common, and it’s fatal for AI, because a model that can’t see yesterday can’t learn anything about tomorrow. 

The first thing the foundation did was remember. Before any forecast, before any agent, the team built a place where the business’s own history accumulated: what was ordered, what shipped, what the plan said versus what happened. Alongside that history came a connector into the main system of record, a way to test outputs against reality, a handful of reusable patterns for how AI intakes data and asks a human for approval, and a written record of every architectural decision and why it was made. 

None of those was the deliverable. The deliverable was a way for their users to interact with the data that was locked behind unavailable systems to the business unit. When the second use case came along, completely unrelated to this specific function, it needed almost none of that built again. When the business wanted to expand into another region, the only thing that changed was the connector. Each use case shipped faster than the last. That compounding of information is what the foundation is for, and it’s the same story we told in the second article about pilots that leave something behind. This is what “behind” looks like. 

Governance that doesn’t mean slow 

One more thing the foundation does, and it’s the one that makes the rest survivable: it lets you put the controls where the consequences are, instead of everywhere. 

Not every AI decision needs a human gate. An assistant that drafts an internal summary needs a basic check and nothing more. A model that changes a price, approves a stage in your workflow or touches a customer, needs an independent second opinion, ideally from a different system than the one that produced the answer, and a named person who signs off. The gates go at the high-consequence moments, not spread evenly across every step so that everything moves at the speed of the slowest approval. 

That’s what most AI governance gets wrong. It treats every use the same, which means either everything is slow or nothing is checked. Matching oversight to risk is the only approach that is sustainable, and you can only pull it off if the foundation exists. You need the register to know what’s running, the one door to see it and manage the costs, and the evaluation framework to judge its performance. 

The muscle to build 

If you intend to run many AI systems, whether you build them or buy them, the capability your organization needs to develop isn’t building more agents. Agents are getting easier to build. The capability you must develop to scale AI usage is the foundation: knowing what you’re running, knowing what it costs, knowing whether it works, and knowing who decides. Evaluation, governance, and the foundation they run on. 

That’s the muscle we’ve learned to prioritize before anything scales — it’s the difference between compounding every new AI effort and starting from zero each time. 

Questions to take back 

Request an AI Briefing

AI has already changed how your people work. Request an AI Briefing with Robots & Pencils to see exactly where your organization sits on the productivity paradox, and what to fix first, across the people, the organization, and the foundation.


About the Author

Brendan Flynn is SVP, Strategist at Robots & Pencils where he heads industry strategy within the Generative & Agentic AI Studio.


Key Takeaways

FAQs

What does it mean to own the foundation instead of the model?

AI models are becoming commodities as leading providers converge in capability, so switching between them is getting easier. What a company builds around the model, its business definitions, its governance, its way of testing outcomes, is the part that keeps its value no matter which model runs underneath it.

What are the five parts of an AI foundation?

A shared definition of core business terms, a single point every AI request goes through, a register of every AI tool with a named owner, a repeatable way to test whether a tool is working, and an approval process sized to the risk of what the tool touches.

Does a company need a foundation if it only buys AI tools and does not build any?

Yes. Governance that only covers tools a company builds misses most of what it actually runs, since most AI tools in a typical company are bought, not built. A foundation gives a company a way to see cost, usage, and performance across every tool, regardless of where it came from.

Does owning the foundation mean centralizing every AI decision?

No. The foundation centralizes standards, shared definitions, and the way tools are tested and tracked. Decisions about what to build, what to buy, and when it goes live in a specific business unit stay local to that unit.

What is the payoff of having a foundation in place?

A company with a foundation can add a new AI tool and know what it costs and whether it works within about two weeks. Without one, every new tool starts from zero and stays an island, disconnected from everything the company has already built.

Higher Education has an AI Governance Blind Spot, and It’s Happening in Every Classroom. 

Robots & Pencils’ new three-part series traces the fissure between the AI policies campuses write and the practices underway in the classroom, then names the standard to close it. 

Faculty are abandoning campus AI bans on their own, one syllabus at a time, because the bans do not match how they want to use AI in their discipline. At the same time, students are living with AI-shaped grades, advising nudges, and early-alert flags they were never told were happening. Robots & Pencils, an applied AI engineering partner known for high-velocity delivery and measurable business outcomes, today published “Classroom AI Governance,” a three-part, data-driven series that names both patterns and proposes the standard that closes the gap between them. 


Start reading Part 1 now: “Classroom AI Governance: The Detection Default” 

The full series, a 10-minute read, is available now. Each article grounds every claim in named, recent higher-education research rather than opinion. 

Campus AI Policy Runs on Two Disparate Documents 

The series is written for the leaders shaping campus AI policy, from academic affairs to IT, and for the faculty and administrators living with the consequences day to day. Most institutions are working from exactly two documents right now: a campus-wide ban with a detection-tool footnote, and, where it exists at all, a scattered set of IT or Registrar guidance no student ever sees. Neither document accounts for the other, and neither is written by the people standing in the room. 

Classroom AI Governance: Faculty Policy and Student Transparency are One Problem 

Classroom AI governance treats that gap as one problem instead of two, argues series author Lindsay Pineda, Senior Delivery Manager for Education at Robots & Pencils. “Most commentary on AI in education treats faculty policy and student-facing transparency as separate stories, one a pedagogy question and the other a compliance question. This series demonstrates they are the same story, told from two sides of one classroom.” 

Pineda’s research names the pattern from inside the classroom. Jason Lacy, Client Partner, Education at Robots & Pencils, hears the same pattern from the leaders funding governance decisions, and keeps bringing the conversation back to one point: “AI governance shouldn’t be measured by how well institutions enforce policy. It should be measured by whether faculty can teach effectively, students trust the experience, and learning improves. That’s ultimately what higher education exists to do.” 

As institutions move from AI experimentation to enterprise adoption, classroom governance will become one of the earliest indicators of whether AI can be scaled responsibly across the institution. Education leaders ready to design AI governance that faculty trust, students understand, and institutions can confidently scale can request an AI Briefing with Robots & Pencils. 

Part 1: Classroom AI Governance – The Detection Default 

Why faculty don’t need another AI policy, they need agency 

This article is part of a three-part series examining how AI is reshaping trust between faculty, students, and the institutions governing them. Reading the full series is recommended.  

Part 2: The Silent Loop | Part 3: The Collision Point 

Elena Marsh writes her syllabus every August at the same kitchen table, and every August for the last three years it has become harder to write. She teaches first-year composition at a mid-size public university, the kind of course where the real subject isn’t grammar, it’s teaching eighteen-year-olds to think in sentences. This year, two days before classes start, the provost’s office sent out the revised campus AI policy. It was a one paragraph, campus-wide notice banning “unauthorized use of generative AI on any graded assignment,” with a footnote instructing faculty to run all submitted essays through the university’s new AI-detection add-on before grading. 

The policy didn’t match anything Elena was trying to do. She wanted her students using AI, but as a brainstorming partner, a sentence-level sparring opponent, a way to see three versions of an argument before picking one, and, most importantly, disclosing to her that they’d used it. The policy, read literally, would flag exactly that workflow as a violation. It said nothing about disclosure. It said nothing about her discipline. It said nothing about the fact that a colleague from two buildings over was teaching a coding bootcamp-style intro course and wanted her students to use AI on every assignment, because knowing how to work with it was the skill being taught. 

So, Elena did what faculty have been doing under the radar for three years now. She wrote her own policy into the syllabus, in language careful enough not to contradict the campus policy outright and hoped no one asked her to reconcile the two.  

The Detection Default 

Call it the detection default, when an institution doesn’t know what else to do about AI, it reflexively reaches for a ban and a detector. It is the easiest policy to write, the easiest to defend to a board of trustees, and the least useful to the person who actually has to run a classroom. It treats faculty as the last line of defense, rather than as the professionals best positioned to decide how AI belongs, or doesn’t, in their own discipline. 

This failure mode mirrors the one The Institutional Intelligence Crisis documented on the operations side of the university, where a single mandated tool or a blanket workaround stripped staff of the judgment that made their work reliable in the first place. In the classroom, the mechanism is identical, and the stakes are just as personal. A philosophy seminar and a coding bootcamp course need different answers to “how should AI be used here,” and the person qualified to set that answer is standing in the room and shouldn’t be constrained to an administrative “one size fits all” approach. 

What the Data Actually Shows 

The instinct to ban is losing ground because faculty are already abandoning it on their own, not because institutions have found something better to replace it with. 

A UC Berkeley study looked at 31,692 course syllabi collected between 2021 and 2025 (Chirikov, reported in Inside Higher Ed, Feb. 2026). It found that academic-integrity concerns, the reason most often given for restricting AI, showed up as the stated rationale in 63% of syllabi in spring 2023, but in only 49% by autumn 2025. 

In place of that blanket justification, faculty are writing rules that vary by task. For example, AI is barred for drafting or revising in 79% of syllabi and for reasoning or problem-solving in 65%, but only 20% ban it for coding tasks, and just 17% for editing or proofreading. Meanwhile, requirements to disclose what AI was used and how jumped from 1% of syllabi to 29% over the same period. 

Faculty are done waiting for institutions to hand down permission to make these distinctions. They’re making them anyway, one syllabus at a time, without a shared template or any institutional backing. 

Meanwhile, the tools institutions lean on to enforce the old model keep failing in predictable, well-documented ways. Independent benchmarking has found that AI-text detectors lose much of their accuracy once a student does any manual editing or paraphrasing, performance that holds up in a controlled test collapses under exactly the kind of light editing real students do (RAID benchmark, Dugan et al., ACL 2024). And detectors don’t fail evenly.  

A Stanford team ran 91 TOEFL essays by Chinese test-takers through seven AI detectors. On average the tools flagged 61.22% of them as AI-generated, while essays from U.S. students came back almost perfectly clean (Liang et al., Patterns, 2023). The reason is mechanical. Detectors score how predictable a piece of writing is, and a student working in a second language under exam pressure reaches for familiar words and safe sentence shapes, which is the exact pattern the tools read as machine-written. 

ETS, the company that owns the TOEFL, took the problem seriously enough to spend a paper on it. Working with 85,567 essays, its researchers tested three fixes: balancing the training data, stripping out the features that track a writer’s language background, and moving the detection threshold. Each one reduced the bias to some degree without gutting accuracy (Jiang et al., 2024). Reduced, not removed. 

And the pressure isn’t easing. The Digital Education Council’s 2026 Global AI in Higher Education Survey, 45,398 responses from students and faculty across 35 countries, found that 88% of students and 77% of faculty now use AI in their coursework or teaching (Digital Education Council, 2026). The 2025 EDUCAUSE AI Landscape Study, meanwhile, found that teaching and learning is now the institutional function most focused on AI adoption, and that faculty training is the single most common element of institutional AI strategic plans (EDUCAUSE , 2025). Everyone agrees training matters and almost no one has funded it at the pace adoption demands. 

Why the Ban Persists Anyway 

If the data so clearly favors discipline-specific judgment over blanket policy, why do so many institutions still default to the ban?  Speed, mostly, and the fact that it’s defensible in a meeting. It also doesn’t require trusting thousands of individual faculty members to make thousands of individual calls. Most institutions didn’t adopt blanket AI policies because they believed them the best pedagogical answer; they adopted them because they were the fastest governance response to a rapidly changing technology. The problem is that what works as an emergency response rarely becomes a sustainable long-term strategy. 

The ban buys speed and legal cover at the price of faculty judgment, and it plays out the same way every time. The first semester, a ban feels like caution. The second semester, with the ban unrevised and unenforceable, it starts to feel like avoidance. By the third semester, most students have routed around it, most faculty have stopped enforcing it as written, and the only thing the policy has reliably measured is the widening gap between what the syllabus says and what happens in the room. 

Underneath that gap is a category error; treating consistency and fairness as the same thing. A single campus-wide rule feels fair because it applies equally to everyone. But applying an identical AI policy to a philosophy seminar building an argument from scratch and a data science studio where AI-assisted coding is the professional standard doesn’t produce fairness, it produces a rule that’s wrong for at least one of them, and often both. Real fairness in this context looks more like a shared space with room for discipline-specific needs. Every student can count on knowing what’s expected of them, even as the specifics vary by course. 

What This Costs an Institution 

At their core, blanket classroom AI policies are retention and liability problems wearing a pedagogy costume. 

On the faculty side, surveys through 2026 have repeatedly found that most instructors receive no formal AI guidance at all, and that the resulting ambiguity is a measurable contributor to instructor burnout. On the institutional-risk side, using an unreliable detector as grounds for an academic-integrity referral is an exposure risk, a due-process problem waiting for a mistaken accusation to surface publicly. On the trust side, every time a policy visibly fails to match classroom reality, it teaches students that the rules are theater, and the real rules are whatever their individual professor decides to enforce. That’s a lesson that risks generalizing beyond AI policy. 

Agency, With a Floor Under It 

The fix isn’t “let every instructor do whatever they want” any more than it’s “one rule for everyone.” It’s giving faculty real authority to set discipline-specific AI policy, backed by institutional infrastructure that makes exercising that authority fast and easy instead of exhausting and tedious. 

In practice, that means a few concrete things. A policy template faculty can adapt to their own course in under an hour, with discipline-specific exemplars, such as what a reasonable AI policy looks like in a lab science, a language course, a studio art class, so no one is solving this from scratch. A fast, low-friction path to revise that policy every term as the tools and the norms shift under everyone’s feet, and real investment in the faculty training that EDUCAUSE’s own data says every institution already claims to prioritize. 

None of that means abandoning consistency altogether. Students should be able to count on a baseline of clarity across every course on their schedule, even when the specific rules differ by instructor and discipline. They need to understand what to expect, not necessarily what’s allowed. That floor is what makes room for faculty agency without turning the whole institution into hordes of disconnected experiments. 

Faculty agency is one half of what happens in that classroom. The other half belongs to the student sitting across from Elena’s desk, who has no idea whether the essay she’s about to turn in will be read by a person, scored by a machine, or some blend of both, and no clear way to find out. 

We explore more of that in Part 2 of this series. Read it now.  

Punch List 

Talk to Robots & Pencils about designing agentic AI for education. Request an AI Briefing.

Note: Elena Marsh is a composite drawn from patterns documented across the sources above, not a named individual or institution. 

About the Author 

Lindsay Pineda is a Senior Delivery Manager at Robots & Pencils, where she leads delivery of an AI-powered student intervention platform for a major public research university. With over 20 years of experience spanning higher education, educational technology, and program and delivery management, she has held leadership roles at a range of organizations across the higher education and edtech sectors. Lindsay spent nearly a decade as an adjunct graduate faculty member at a large online university facilitating master’s level courses in project management leadership and PMP exam preparation while contributing to curriculum and instructional design. A PMP-certified leader with master’s degrees in psychology and management, she brings a rare blend of strategic delivery expertise and firsthand experience in online course facilitation and the student learning experience. 


FAQs

Q: What is “the detection default”?
A: The reflexive move institutions make when they don’t know how to handle AI: ban it, then run submissions through a detection tool. It’s the fastest policy to write and the least useful to the person running the classroom.

Q: Are faculty actually following campus AI bans?
A: No, and the data shows it. A UC Berkeley study of 31,692 syllabi found integrity concerns as the stated rationale for restricting AI dropped from 63% (spring 2023) to 49% (autumn 2025), while disclosure requirements jumped from 1% to 29% over the same period.

Q: Do AI detectors actually work?
A: Not reliably. The RAID benchmark found detector accuracy collapses once a student does any manual editing. Worse, a Stanford study found detectors flagged 61.22% of TOEFL essays from Chinese test-takers as AI-generated versus near-zero for U.S. students, because detectors penalize predictable phrasing, a pattern common in second-language writing.

Q: What’s the alternative to a blanket ban?
A: Discipline-specific policy, set by faculty, with an institutional floor: a fast policy template every instructor can adapt, exemplars by discipline, real training investment, and a disclosure baseline students can count on regardless of course.

Q: What does this cost an institution that keeps the ban?
A: Faculty burnout from ambiguous guidance, legal exposure from using unreliable detectors as sole grounds for integrity referrals, and erosion of student trust when the written policy doesn’t match classroom reality.

Q: How does this connect to the rest of the series?
A: Part 1 covers faculty agency. Part 2 (The Silent Loop) covers the student side, undisclosed AI in grading and advising. Part 3 (The Collision Point) unifies both into one governance standard.


Key Takeaways

Part 2: Classroom AI Governance – The Silent Loop 

Earning trust when AI is making decisions about you 

This article is part of a three-part series examining how AI is reshaping trust between faculty, students, and the institutions governing them. Reading the full series is recommended.  

Part 1: The Detection Default | Part 3: The Collision Point 

Maya Chen is a first-year student in Elena Marsh’s composition course, and two weeks into the semester she has already sensed a contradiction she can’t quite put into words. The program she is enrolled in requires every first-year writing student to run their essay outline through the university’s approved AI tool before drafting; not because Elena asked for it, but because the college wants to show it’s “building AI literacy.” Maya does what the assignment asks. She types a prompt into the tool, screenshots the output for the completion credit, and writes her actual essay the way she always has. She has learned nothing about the tool, but then, that was never really the point of the exercise. 

Then she submits her final draft, and it passes through the university’s writing-assessment pipeline, which is mix of an AI-assisted feedback tool and Elena’s own read. It comes back marked down half a grade for “irregular phrasing,” with no further explanation. Maya sits with that for a while. She wasn’t allowed to use AI to help write the essay; using it that way would have been an integrity violation. But something used AI to help judge it. She has no way of knowing which parts of her essay created the flag, whether a person looked closely at those sentences before the grade was finalized, or who she’d even ask to get more clarification. The tool that graded her didn’t have to explain itself, but she is required to do so in every essay. It just does not feel like a fair trade to Maya. 

The Silent Loop 

Call it the silent loop. The growing set of decisions that touch a student’s academic life; a grade, a feedback comment, a flag on their advising file, a nudge to switch majors, that are increasingly made or shaped by AI, all without the student being told when or how. No one set out to hide anything. Disclosure was never built into the system in the first place, and it’s nobody’s job is to notice the gap. 

 The loop is already running in places most students never see. Early-alert systems score retention risk. Advising platforms surface nudges. Assessment tools flag phrasing. Most of these are genuinely well-intentioned programs, and some of them demonstrably work. But in April 2026, the National Student Legal Defense Network published a Student AI Bill of Rights whose first article asserts that students have a right to know “when, where, and how AI systems are being used to evaluate them, track them or make decisions about their educational future” (National Student Legal Defense Network, 2026). Nobody writes that sentence unless the current answer is no. The tools work; whether the student knows they’re running is a separate question, and mostly an unasked one. 

What Students Already Know… and What They Don’t 

Here’s the part that should recalibrate how institutions think about this; students are not the passive party in the AI story. A 2026 Digital Education Council survey that found 77% of faculty now use AI in teaching, also found 88% of students already using it in their own learning (Digital Education Council, 2026). This generation does not need to be introduced to technology. They are more fluent in it than most of the adults setting policy for them. Students know this, and that adds fuel to the fire. Maya isn’t confused about what AI can do, she’s frustrated that the institution gets to use it without the same disclosure it demands from her. 

Research on AI-assisted grading backs up the instinct behind that frustration. A small study looked at 27 undergraduate computer science students grading a programming project, not an essay, but it asked the same underlying question: do students trust AI feedback as much as they trust a human’s? Even when the AI’s scores and clarity ratings matched or exceeded a human teaching assistant’s, most students still preferred the human. Sixty percent of students rated the TA’s feedback as fairer, and 55% said they trusted it more overall (Riahi, Storozhevykh & Catete, 2026). They chose the grader who gave them worse marks and murkier explanations. 

Students Want the Why 

Students consistently pointed to the same specific frustration Maya has; the AI could tell them what was wrong, but not why it mattered, or what to do next. This is the kind of contextual judgment that comes from an instructor who knows where a particular student is in their development. That same complaint showed up, almost word for word, in an account heard directly while researching this piece. A parent said her daughter’s high school teacher admitted that an AI tool couldn’t grade this particular ninth-grader accurately; it couldn’t tell what the student already knew versus what she still didn’t. It could grade the words on the page, but not the student behind them. 

Jisc’s 2025 survey of student perceptions of AI found the same pattern. Students want clear institutional guidance on AI use and consistently say they value personalized, human feedback over automated alternatives. This is not because the automation is inaccurate, but because it can’t yet account for who they specifically are (Jisc, 2025). 

Why the Trust Gap Is Actually a Retention Problem 

It would be easy to file this under ethics and move on, but that undersells what’s actually at risk. Students who feel monitored or misjudged by systems they don’t fully understand rarely file a complaint; they disengage instead. Students are not measuring an institution’s AI governance against another university’s policy manual. They’re measuring it against every other digital experience they have every day, from their banking app to their streaming service. Held to that standard, an AI decision that feels opaque, inconsistent, or impossible to question doesn’t just cost the tool their confidence; it costs the institution behind it. 

The same can be said for faculty not enforcing bans they don’t believe in. A flagged essay with no explanation costs Maya half a letter grade, and it teaches her that the system’s judgments about her are unappealable. This lesson generalizes fast to other areas such as advising nudges, degree-progress flags, and every other place AI touches her file. Gen Z and Gen Alpha students are, by every available measure, more AI-literate and more skeptical of unclear automated decisions than most institutional AI rollouts assume. An institution that treats that skepticism as a communications problem rather than a design problem will keep manufacturing the mistrust it’s trying to avoid. 

What Human Agency Actually Requires 

“Human-centered AI” has become a phrase institutions attach to almost anything. For a student, it means three concrete things. 

The first is disclosure. A student should always be able to find out, without having to ask directly, when AI shaped a decision about them, such as a grade, a flag, a recommendation. The second is a working path to a human. Not a hidden one, not one that requires escalating through multiple offices, but an accessible route to someone with the authority to actually look again. The third is explainability calibrated to the stakes. A scheduling suggestion doesn’t need the same depth of explanation as a probation flag or a grade that affects a scholarship. Treating every AI touchpoint with the same disclosure process either buries the important ones in noise or makes the whole system too heavy to use. The standard should scale with what the student stands to lose. 

The first two of those are already written down. The Student AI Bill of Rights asks for disclosure, and it asks that “automated systems should not be the final arbiter of high-stakes decisions affecting a student’s admission, academic standing, financial stability or other aspects of fundamental well-being” (National Student Legal Defense Network, 2026). Which is to say the standard being proposed here is not a radical one, and institutions will not get to claim they were never told. 

None of this asks institutions to slow down AI adoption, rather it asks them to build the disclosure and recourse in from the start. This is generally the way a well-governed system is designed with an audit trail from day one rather than bolted on after something goes wrong. 

From far away, it can look like teachers wanting control over their own work and students wanting to be trusted are two completely separate issues; they aren’t. The exact moment Elena Marsh’s grading tool flags Maya’s essay is the same moment both stories collide; one instructor exercising legitimate professional judgment, one student on the receiving end of a decision she never saw coming.

That collision is discussed in Part 3 of this series. Read it now. 

Punch List 

Note: Maya Chen is a composite illustration informed by parent- and educator-reported experience with mandated AI tools and AI-assisted grading, not a named individual. The high school teacher’s account referenced above is a first-hand anecdote relayed to us during research for this piece, not a published or independently verified source — it’s included as illustrative color, not as data. 

Talk to Robots & Pencils about designing agentic AI for education. Request an AI Briefing. 

Lindsay Pineda is a Senior Delivery Manager at Robots & Pencils, where she leads delivery of an AI-powered student intervention platform for a major public research university. With over 20 years of experience spanning higher education, educational technology, and program and delivery management, she has held leadership roles at a range of organizations across the higher education and edtech sectors. Lindsay spent nearly a decade as an adjunct graduate faculty member at a large online university facilitating master’s level courses in project management leadership and PMP exam preparation while contributing to curriculum and instructional design. A PMP-certified leader with master’s degrees in psychology and management, she brings a rare blend of strategic delivery expertise and firsthand experience in online course facilitation and the student learning experience. 


FAQs

Q: What is “the silent loop”?
A: The growing set of decisions touching a student’s academic life, a grade, a feedback flag, an advising nudge, a major-change suggestion, that AI increasingly shapes without the student being told when or how it happened.

Q: Do students actually trust AI-generated feedback?
A: Not as much as human feedback, even when it’s just as good. A study of 27 undergraduate computer science students found that even when AI feedback matched or exceeded a human TA’s accuracy, 60% still rated the TA’s feedback as fairer and 55% trusted it more (Riahi, Storozhevykh & Catete, 2026).

Q: Isn’t this generation comfortable with AI making decisions about them?
A: They’re comfortable using AI, not comfortable with asymmetry. 88% of students already use AI in their own learning (Digital Education Council, 2026), which makes them more attuned to, not less bothered by, an institution using AI on them without the same disclosure it demands from them.

Q: What does the Student AI Bill of Rights actually require?
A: Published by the National Student Legal Defense Network in April 2026, its first article states students have a right to know when AI is evaluating, tracking, or deciding their educational future, and that automated systems shouldn’t be the final arbiter of high-stakes decisions.

Q: Why treat this as a retention issue instead of an ethics issue?
A: Students don’t file complaints when they feel misjudged by an opaque system, they disengage. They measure institutional AI against their banking app or streaming service, not against a policy manual, and an unexplainable decision costs the institution their confidence.

Q: What are the three things human agency actually requires?
A: Disclosure (knowing when AI shaped a decision), a working path to a human who can look again, and explainability calibrated to stakes, a scheduling nudge needs less explanation than a probation flag or scholarship-affecting grade.


Key Takeaways

Part 3: Classroom AI Governance – The Collision Point 

Where faculty agency meets student trust 

his article is part of a three-part series examining how AI is reshaping trust between faculty, students, and the institutions governing them. Reading the full series is recommended.  

Part 1: The Detection Default | Part 2: The Silent Loop 

Here is the moment, the collision point, stated plainly. Elena Marsh, exercising the exact discipline-specific judgment Part 1 argued she should have, uses an AI-assisted feedback tool to help her manage grading load across 90 first-year essays. It’s a reasonable professional choice, made in good faith, inside a policy her department endorses. Maya Chen, sitting on the other side of that same decision, receives a grade shaped in part by a tool she was never told was involved and with no path to ask why. Both things are true in the same instant. Elena is exercising legitimate pedagogical agency. Maya is a student having a decision made about her by a system she can’t see. Neither of them did anything wrong. The institution simply never designed for the fact that its faculty-agency policy and its student-trust policy would collide in this room, over this piece of work. 

Two Workstreams, One Collision Point 

Most institutions treat faculty AI policy and student-facing AI transparency as two different projects, run by two different offices, and on two different timelines. Faculty policy usually lives with the Provost or a Center for Teaching and Learning. Student-facing disclosure, when it exists at all, tends to live with IT, the Registrar, or student affairs. It often doesn’t exist as a formal policy so much as an assumption that someone else is handling it. Each office can point to real progress on its own workstream. Neither has been asked to think about the moment those two workstreams meet; the instant an instructor’s tool becomes a student’s outcome. 

This is the same structural blind spot The Institutional Intelligence Crisis found across university operations: departments run independently, and no one is responsible for what falls through the cracks between them. In the classroom, that crack isn’t between departments; it’s between the person given the power and the person affected by how they use it. Almost no institution has a governance table where both of them sit. 

Why the Fix Isn’t Another Committee Handing Down Rules 

The instinct, once an institution notices this gap, is to convene an IT-and-Provost governance committee and issue joint guidance. That instinct reproduces the exact failure both prior pieces in this series documented; policy written by the people furthest from the room, applied to the people standing in it. A governance model built to hold faculty agency and student trust together must include faculty senate representation, because faculty are the ones who must live inside whatever gets decided, and it must include actual student voices. And not just a single student representative added to satisfy an optics requirement, because students are the ones the decisions land on. 

This is structurally different from the accountability-owner model that works for administrative AI. A single named owner for a workflow tool makes sense when the tool serves one office and one function. That structure doesn’t work here, because the classroom isn’t one function; it’s two people with different relationships to the same decision. A governance structure that represents only one of them will keep producing policy that only looks complete but functions incompletely. 

A Standard Simple Enough to Actually Adopt 

The practical version of this doesn’t need to be complicated, and it shouldn’t wait for a perfect governance model to be built before any course adopts it. Any instructor, in any discipline, can commit to two things without needing campus-wide uniformity on how AI gets used. One, tell students what AI was used for on a given piece of work, and two, give them a real, findable way to ask for a second look if they think it got something wrong. 

That’s the whole standard. It doesn’t require Elena to disclose her exact tool stack or her grading workflow in granular detail. It doesn’t require the registrar to build a new system before anyone can use it. It requires the two things students in Part 2 said they wanted; to know how the decision was reached (disclosure), and to have somewhere to go to ask questions (recourse). Institutions already building agentic systems with real governance, which includes identity, oversight, and an audit trail designed in from the first sprint rather than bolted on after a trust failure, tend to treat this kind of disclose-and-recourse checkpoint as a basic architectural requirement. Classroom AI deserves the same standard the best-engineered institutional systems already hold themselves to. 

Scaled up, that same logic becomes the governance table’s actual job. Which is not dictating how every course uses AI, but making sure every course, regardless of how it uses AI, meets that standard. 

Designing the Relationship, Not Just the System 

The thread running through all three pieces in this series is the same; human-centered agentic AI in higher education is not primarily a data architecture problem, and it’s not primarily an operations problem; it’s both. Both of those are real, and both are already being worked on elsewhere. It’s a relationship problem, between two people who are physically in the same room and structurally treated as if they’re solving two unrelated problems. 

ASU’s framing for its Agentic AI and the Student Experience summit this October puts it well; the goal is designing AI systems that “enhance human agency, expand access, and strengthen learning in meaningful ways” (ASU, 2026). That framing only works if “human agency” means both humans in the room; the instructor deciding how AI belongs in her discipline, and the student who deserves to know when it’s being used on her. Institutions that get this right will avoid a trust problem, and they’ll have a classroom-level foundation solid enough to make everything already being built at the operational layer worth scaling. 

The institutions that solve this well won’t simply have better AI governance. They’ll strengthen one of the most important relationships on campus: the trust between faculty, students, and the institution itself. That trust becomes the foundation for every future AI initiative. 

As institutions move from AI experimentation to enterprise adoption, classroom governance will become one of the earliest indicators of whether AI can be scaled responsibly across the institution. Education leaders ready to design AI governance that faculty trust, students understand, and institutions can confidently scale can request an AI Briefing with Robots & Pencils.

Punch List 

Note: Elena Marsh and Maya Chen are composite illustrations carried through from Parts 1 and 2, not named individuals. 

Lindsay Pineda is a Senior Delivery Manager at Robots & Pencils, where she leads delivery of an AI-powered student intervention platform for a major public research university. With over 20 years of experience spanning higher education, educational technology, and program and delivery management, she has held leadership roles at a range of organizations across the higher education and edtech sectors. Lindsay spent nearly a decade as an adjunct graduate faculty member at a large online university facilitating master’s level courses in project management leadership and PMP exam preparation while contributing to curriculum and instructional design. A PMP-certified leader with master’s degrees in psychology and management, she brings a rare blend of strategic delivery expertise and firsthand experience in online course facilitation and the student learning experience. 


FAQs

Q: What is “the collision point”?
A: The exact moment a faculty member’s legitimate AI-assisted grading choice becomes a student’s outcome, without the student ever knowing a tool was involved or having a way to ask why. Faculty agency and student disclosure aren’t separate problems, they meet in the same room, over the same piece of work.

Q: Why can’t a joint IT-and-Provost committee just fix this?
A: Because that reproduces the exact failure Parts 1 and 2 documented, policy written by people furthest from the classroom, applied to the people standing in it. A governance table needs faculty senate representation and real student voices, not one token student seat.

Q: What’s the actual two-part standard being proposed?
A: Any instructor, in any discipline, can commit to two things without campus-wide uniformity: tell students what AI was used for on a given piece of work, and give them a real, findable way to ask for a second look.

Q: Does this require a new system or registrar build-out?
A: No. It doesn’t require disclosing a full tool stack or grading workflow in detail, and it doesn’t require IT to build anything new before an instructor can adopt it. It’s disclosure plus recourse, nothing more.

Q: Is this a data problem or a relationship problem?
A: Both are real, but the series argues it’s primarily a relationship problem, between two people physically in the same room who are structurally treated as if they’re solving unrelated problems.

Q: How does this connect to institutional AI governance generally?
A: The same disclose-and-recourse checkpoint that well-engineered agentic systems already build in from the first sprint (identity, oversight, audit trail) should apply to classroom AI. Getting this right becomes the foundation for scaling every other AI initiative on campus.


Key Takeaways

Designing Voice AI for Truck Drivers: The Human Questions Behind E.L.L.A. 

When FleetHub.AI set out to build a voice-first co-pilot for commercial fleets, the design challenge had little to do with screens and everything to do with judgment.

The machines got smarter. So did the human work.

There were no mockups. No component libraries. No color tokens. No screens to stare at in a design review while someone asks if the font could be slightly bigger.

Just a different kind of question than I was used to asking:

Who is E.L.L.A., and what does this actually do?

I am part of the Robots & Pencils team working with FleetHub.AI to shape E.L.L.A.’s experience, and I quickly realized we weren’t designing another interface. We were designing how an intelligent system should behave around a person.

Inside the Design of E.L.L.A., FleetHub.AI’s Voice Co-Pilot for Truck Drivers

FleetHub.AI is building a connected technology stack for commercial trucking fleets. At the center of that vision is E.L.L.A., a voice-first AI co-pilot designed to live in the cab alongside truck drivers, not as another app competing for attention, but as an intelligent layer connecting the systems already at work around them. Underneath, E.L.L.A. runs on Amazon Nova 2 Sonic for real-time voice and Amazon Bedrock for reasoning, with AWS Lambda, Amazon API Gateway, and Amazon DynamoDB handling everything in between.

We weren’t designing another interface for a driver to look at. We were designing how an intelligent system should listen, interpret, respond, and know when to stay out of the way.

E.L.L.A. had to work across a complicated technology environment while making the experience feel simple for the person behind the wheel. A co-pilot, not another screen. Something that needed to understand the difference between a driver who needed a heads-up about road conditions and a driver who needed silence. Something that had to know when to surface a delivery conflict, when to help, when to act, and when to simply let the road go by.

Eyes on the road. Hands on the wheel.

In this environment, every interaction that asks a driver to look away carries a cost. Visual design wasn’t irrelevant, but it couldn’t be the primary way E.L.L.A. communicated. The experience had to work without requiring the driver’s eyes.

Designing a Voice AI Persona for Trucking Means Asking Human Questions, Not UX Questions

The difference this time was what I was actually designing.

Not an experience built around a person. A person.

Not a screen a driver would glance at, but a presence a driver would talk to.

That meant designing a character: a voice, a sensibility, a set of behaviors, and a sense of when to speak and when to pause. It was one of the most human design problems I’d worked on, and it had almost nothing to do with visual design.

What would I want to hear if I’d been driving for six hours? What information would feel helpful at mile three and overwhelming at mile forty? If I was running late and stressed, would I want a conversation, or would I want E.L.L.A. to handle it quietly and tell me when it was done?

At what point does a helpful nudge become noise?

And once you answer that, what’s the next layer?

What if the driver just got off a difficult call? What if traffic is about to add forty minutes to an already-long day? What if something changes in the delivery workflow while the driver is moving?

These weren’t UX questions in the way I’d always framed them.

They were human questions.

The more I worked on E.L.L.A., the less I thought of her as a voice interface and the more I realized we were designing how she should behave — how she should understand context, make judgments, respond to a person, and know when to speak or stay quiet.

That distinction mattered.

A traditional interface gives someone a place to go and things to interact with. E.L.L.A. had to do more of the work. Her behavior had to account for what was happening around the driver, determine what mattered, and decide how — or whether — to bring it into the conversation.

That meant we weren’t just designing what E.L.L.A. would say.

We were designing the conditions under which she should speak at all.

Answering those questions required me to sit in the driver’s seat, not literally, but fully. To model a person, not just a user flow.

That project changed something in how I see our work.

E.L.L.A. made one thing impossible to ignore: when you strip away the screen, design doesn’t disappear. It moves somewhere else.

It moves into judgment.

Why Voice AI Design for Trucking Demands Human Judgment

We talk a lot right now about how AI is transforming design. The tools are faster. The cycles are compressed. What used to take three weeks of exploration can surface in three hours. I won’t argue with any of that.

But I think we’re sometimes tempted to let the speed of the tools define the conversation. To make the story about efficiency. About throughput. About how many design variations we can generate before lunch.

And I’d push back on that, not because the speed isn’t real, but because it’s not the point.

E.L.L.A. didn’t need more variations. She needed judgment.

She needed someone to think carefully about what a human being, inside a specific context, on a specific kind of day, would actually find useful.

No tool was going to figure that out.

That thinking had to come from somewhere, from someone willing to put themselves in those shoes and stay there long enough to feel the weight of the question.

And in trucking, that context matters.

A driver isn’t sitting at a desk waiting for information. They’re managing a long day on the road, navigating traffic, schedules, deliveries, communications, compliance requirements, and everything else that comes with keeping a truck moving.

The challenge isn’t necessarily a lack of information.

It’s too much of it, arriving through too many systems, at too many moments.

Good voice AI has to make that complexity feel simpler, not add another layer of noise.

That’s where judgment becomes part of the design.

When the Interface Becomes Behavior

Here’s what I keep coming back to.

With traditional digital products, we spend a lot of time designing what people see, where they tap, what happens next, and how information is organized on a screen.

With conversational and agentic AI, some of that work moves upstream.

We’re designing what the system knows. What it notices. What it prioritizes. What it does on someone’s behalf. What it says. How it says it. And, perhaps most importantly, what it chooses not to say.

The design artifact becomes behavior, not interface.

For E.L.L.A., that meant thinking about timing as carefully as content.

The right information at the wrong moment is still the wrong experience.

A useful notification can become an interruption. A well-intentioned prompt can become noise. A perfectly designed response can still be wrong if it arrives when the driver doesn’t need it.

The intelligence isn’t just in having the answer.

It’s in knowing when the answer matters.

The Design Work That Doesn’t Change

Here’s what I keep landing on:

The canvas has changed.

Interfaces are becoming ambient, conversational, and layered. With E.L.L.A., we weren’t just designing a screen or even a voice. We were designing an intelligent layer between a person and the systems surrounding them — one that had to account for what mattered, what could wait, and what didn’t need to be said at all.

We’re no longer just designing what people see.

We’re designing the conditions under which a system decides how to meet them.

That’s a bigger, stranger, more interesting problem than we’ve had before.

But the core of the work?

Unchanged.

We are still, fundamentally, in the business of understanding people well enough to know what they need before they’ve figured out how to ask for it.

That has always been the job.

Screens were just one era’s answer to how you meet a person where they are.

The tools get faster. The outputs get more sophisticated. The systems get more capable.

And every time the canvas shifts — from web to mobile, from mobile to voice, from voice to whatever comes next — the designers who do the best work are the ones who go back to the same question.

Not what can I build? What does this person actually need?

E.L.L.A. made that question feel especially real because the person at the center of the experience can’t always look at what we’ve built.

So, we had to design for everything else.

For attention. For timing. For trust. For context.

For knowing when to speak.

And when to let the road go by.

That’s what good design does.

It always has.

Learn more about Robots & Pencils’ solutions for the transportation industry.

About the Author

Brad Istnick is Principal Experience Creative at Robots & Pencils, where he helps shape how intelligent systems speak, listen, and decide — turning complex AI behavior into interactions that feel less like software and more like presence.

Key Takeaways

FAQs

What is voice AI for truck drivers?

Voice AI for truck drivers is AI technology that allows drivers to interact with information and fleet systems through natural conversation rather than relying primarily on screens, tapping, or manual navigation. The goal is to make information and assistance available while minimizing unnecessary visual and manual interaction with technology.

How does E.L.L.A. work as a voice AI co-pilot for commercial fleets?

E.L.L.A. is FleetHub.AI’s voice-first AI co-pilot, designed specifically for the commercial trucking environment. Rather than functioning as another standalone application, E.L.L.A. is designed to work as an intelligent layer across the systems and information surrounding the driver, helping determine what matters, when it matters, and how it should be communicated.

What is E.L.L.A. built on?

E.L.L.A. runs on Amazon Nova 2 Sonic for real-time voice and Amazon Bedrock, using Claude 3.5 Haiku for reasoning. AWS Lambda, Amazon API Gateway, and Amazon DynamoDB handle everything underneath the conversation.

How is E.L.L.A. different from a traditional voice assistant?

E.L.L.A. isn’t designed as a general-purpose voice assistant. She’s purpose-built for commercial trucking and the specific operational context of a professional driver. Her value comes not simply from answering questions, but from understanding the environment in which the driver is working and helping turn complex fleet information into timely, conversational assistance.

How is designing voice AI different from traditional UX design?

Designing voice AI moves much of the design work beyond screens and visual flows. Designers have to define a system’s voice, personality, behaviors, timing, decision-making, and rules for when it should speak and when it shouldn’t. The design artifact becomes behavior, not just interface.

How can voice AI reduce driver distraction?

Voice AI can reduce the need for drivers to physically interact with multiple applications while driving by making information and certain interactions available conversationally. The goal isn’t to eliminate every visual interface, but to reduce unnecessary moments when a driver needs to look away from the road to interact with technology.

What makes a good voice AI experience for truck drivers?

Context-awareness and judgment. A good voice AI understands that the same information can have very different value depending on what the driver is doing, what has already happened, and what requires attention right now. It earns trust by getting the timing right — not simply by having the right answer.

How FleetHub.AI is Redefining Fleet Intelligence with E.L.L.A.

A Conversation with FleetHub.AI CEO Marc Lonson

FleetHub.AI has been helping transportation and logistics companies simplify fleet operations through connected technology. By bringing together rugged hardware, wireless connectivity, mobile device management (MDM), and mobility solutions into a single ecosystem, FleetHub.AI has focused on one mission: making it easier for fleets to stay connected, productive, and compliant.

Now, the company is taking the next step in that evolution.

Meet E.L.L.A.! FleetHub.AI’s Intelligent Operating Layer. Purpose-built for the transportation industry, E.L.L.A. combines real-time operational data with agentic AI to proactively support drivers, fleet managers, and operations teams. Rather than adding another dashboard or chatbot, E.L.L.A. works behind the scenes, delivering the right information at the right time to help fleets operate more efficiently.

We sat down with FleetHub.AI CEO Marc Lonson, and our own Scott Young to discuss the inspiration behind E.L.L.A., how it builds on FleetHub.AI’s long-term vision, and why intelligent fleet operations are about empowering people, not replacing them.

Question: What inspired you to create E.L.L.A., and why now?

Marc Lonson, CEO Fleethub.AI: FleetHub.AI has always been focused on solving real-world challenges for transportation companies. Every solution we’ve developed has been driven by one question: How can we make life easier for the people who keep freight moving?

I’ve spent many years working alongside fleets and have developed an incredible appreciation for professional drivers. They’re the backbone of our economy, yet they’re expected to navigate increasingly complex technology, regulations, and day-to-day operational demands.

E.L.L.A. grew out of that understanding.

The timing couldn’t be better. Artificial intelligence has matured to the point where it can provide meaningful, contextual assistance, not just answer questions. Combined with the connected platform we’ve built at FleetHub.AI, we now have the ability to proactively support drivers and fleet teams in ways that simply weren’t possible a few years ago.

For us, E.L.L.A. isn’t about introducing AI for the sake of AI. It’s about making every interaction with technology simpler, smarter, and more valuable.

Question: Everyone seems to be adding AI to their products. How is E.L.L.A. different?

Marc Lonson: AI has become a buzzword, and it’s easy to understand why. But simply adding AI doesn’t automatically create value.

What makes E.L.L.A. different is the context.

Because she’s built into the FleetHub.AI ecosystem, E.L.L.A. understands the operational environment she’s working in. Through secure APIs and webhook integrations, she consumes data that’s directly relevant to each fleet, each driver, and each workflow.

That means she isn’t providing generic answers. She’s delivering proactive, real-time guidance that’s specific to what’s happening in that moment.

Whether it’s helping with Hours of Service, inspection readiness, tablet health, or operational workflows, E.L.L.A. quietly works in the background, stepping in only when she’s needed.

That’s why we describe E.L.L.A. as an Intelligent Operating Layer rather than just another AI application.

Question: How does E.L.L.A. fit into the way fleets already operate today?

Marc Lonson: One of our biggest priorities was making sure customers wouldn’t have to change the way they operate.

FleetHub.AI built a connected mobility ecosystem that brings together hardware, connectivity, device management, and operational technology. E.L.L.A. builds on that foundation.

Because she integrates through existing APIs and webhooks, fleets don’t need to replace their current systems or redesign their workflows. Drivers continue using the tools they’re already familiar with, while E.L.L.A. works behind the scenes, connecting information across the platform and providing proactive guidance when it’s most valuable.

Technology should adapt to the customer, not the other way around.

Question: What are the biggest operational challenges E.L.L.A. is designed to solve for fleet operators?

Marc Lonson: Every fleet faces countless operational decisions every day. Many of them are repetitive, time-consuming, and preventable.

Our goal is to remove as much operational friction as possible by giving people the information they need before they have to ask for it.

Today, E.L.L.A. can help fleets by:

Individually, those capabilities save time. Together, they reduce administrative workload, improve compliance, minimize downtime, and help drivers stay focused on what matters most, driving safely and efficiently.

Question: How will fleet managers experience E.L.L.A. differently from the software they use today?

Marc Lonson: Traditional software tells you what happened yesterday. E.L.L.A. is designed to help you avoid tomorrow’s problems.

Fleet managers should notice fewer calls back to the office because drivers already have the information they need. They’ll spend less time assigning drive time after unidentified driving events and less time responding to routine operational questions.

Instead of creating another dashboard full of alerts, E.L.L.A. reduces the number of alerts that need attention in the first place.

That’s a fundamentally different experience.

Question: Is E.L.L.A. replacing people or helping them work smarter?

Marc Lonson: That’s an important question because there’s a lot of uncertainty around AI.

The answer is simple; E.L.L.A. isn’t here to replace people.

She’s here to support them.

Think of E.L.L.A. as an intelligent co-pilot that works alongside drivers, dispatchers, safety managers, and fleet operators. She proactively identifies opportunities to help, surfaces relevant information, and reduces repetitive manual tasks so people can focus on higher-value work.

The expertise, judgment, and relationships that make great fleet professionals successful will always come from people. E.L.L.A.’s role is to amplify those strengths.

Question: If fleet operators remember one thing about E.L.L.A., what do you hope it is?

Marc Lonson: I hope they remember why we built her.

Everything starts with the driver.

When drivers have the information they need exactly when they need it, they’re more confident, more productive, and less distracted. That creates a ripple effect across the entire organization.

Safer drivers lead to more efficient operations. More efficient operations lead to better customer service. Better customer service leads to stronger businesses.

If E.L.L.A. improves the driver’s day, everyone benefits.

Question: FleetHub.AI has always focused on simplifying fleet operations. How does E.L.L.A. build on that mission?

Marc Lonson: FleetHub.AI has never been about selling individual products.

We’ve always been focused on building a connected ecosystem for transportation companies.

Our platform brings together rugged hardware, wireless connectivity, mobile device management, and mobility technology into one integrated experience. E.L.L.A. is the intelligence layer that brings all of those capabilities together.

Instead of simply collecting operational data, FleetHub.AI can now help fleets understand that data, anticipate issues, and take action before those issues affect drivers or operations.

That’s where the real value of AI comes from, not replacing existing technology, but making every part of the platform smarter.

Question: What excites you most about what E.L.L.A. will enable for customers over the next few years?

Marc Lonson: This is only the beginning.

One of the things I’m most excited about is learning alongside our customers. As fleets continue using E.L.L.A., we’ll measure the impact she’s having and listen carefully to the challenges they still want solved.

Transportation never stands still, and neither will FleetHub.AI.

We already have exciting ideas for where E.L.L.A. goes next, but we’ll keep those under wraps for now. What’s important is that we’ve built a platform designed to evolve continuously, helping our customers solve today’s problems while preparing them for tomorrow’s opportunities.

Question: From an engineering perspective, what made E.L.L.A. a unique and exciting project to build?

Scott Young: E.L.L.A. sits in a different category from most “add AI to it” products. Most fleet AI tools report to the back office after something already happened. E.L.L.A. had to work the opposite way: catch a problem before it happens, in a cab, with a driver who can’t stop and stare at a screen. That constraint made this fun to build. Voice-first, proactive, event-driven: every decision had to hold up in the real conditions of a moving truck, not a demo environment. Building an agent that acts on its own judgment inside someone’s workday, without asking them to change how they work, is a genuinely hard problem. That’s exactly the kind we like.

Question: When FleetHub first shared its vision for E.L.L.A., what stood out to your team, and what made you excited to help bring it to life?

Scott Young, EVP of Growth and Strategic Alliances, Robots & Pencils: What stood out from day one was how clearly Marc and Ella understood the real cost of driver friction, drawn from decades spent inside fleet operations, telematics, and device security. They came to us with a clear ask: build a co-pilot that earns a driver’s trust in silence, showing up right when needed and staying out of the way otherwise. That clarity made the partnership work. We knew exactly who we were building for and why it mattered.

Question: Looking back from concept to launch, what are you most proud of, and what makes E.L.L.A. different from other AI solutions entering the market?

Scott Young: What we’re most proud of is that E.L.L.A. was built to work for the driver first, ahead of the back-office dashboard. It acts in the moment: before a violation gets recorded, before hours run out, before an inspection catches a driver off guard. Most AI entering this market is built to inform someone later. E.L.L.A. is built to help someone now, and that timing is the real differentiator. (edited) 

Looking Ahead…

For FleetHub.AI, E.L.L.A. is more than the company’s latest innovation – it’s the next evolution of its mission to simplify fleet operations through intelligent technology.

By combining connected hardware, wireless connectivity, mobile device management, operational data, and agentic AI into a single ecosystem, FleetHub.AI is redefining how transportation companies interact with technology. Instead of asking drivers and fleet managers to adapt to disconnected systems, the platform works together as one intelligent environment, providing proactive guidance, reducing operational complexity, and empowering people to make better decisions every day.

As the transportation industry continues to evolve, so will FleetHub.AI. And with E.L.L.A. leading the way, the future of fleet intelligence isn’t just connected, it’s proactive, contextual, and built around the people who keep the world moving.

Learn more about Robots & Pencils solutions for transportation and logistics companies.

Pt. 2: Repo-Native Delivery – The Operating Model 

Part 2 of 2. Part 1 made the case for moving delivery out of Jira and into the repo. This is the how, concrete enough to copy. 

In Part 1, I argued that a tracker quietly makes humans the integration layer, and that moving delivery artifacts into the repo hands much of that assembly work to the tooling. 

This is the operating model that fell out of that idea. 

The goal is simple: 

Every artifact should be readable by both a human and an assistant without translation. 

Stories, decisions, sprint history, status updates, technical plans, reports: all of it lives in a form that the team can read directly and that an assistant can traverse without APIs, connectors, or synchronization. 

None of this is exotic. It’s markdown files, Git, and an agent-aware editor like Cursor or Claude Code. 

The discipline is what makes it work. 

Two repos, two rhythms 

The first decision is to stop forcing artifacts and code to share a history. 

They have different cadences, different gates, and different owners. 

So they get different repos. 

The delivery repo holds artifacts only: stories, sprints, decisions, context, prompts, and reports. 

The code repos stay exactly as they should be: feature branches, pull requests, CI, and release workflows. 

Each code repo’s CLAUDE.md points back to the delivery repo so that any assistant reading the code also has access to the project’s intent and decisions. 

That separation solves a surprisingly common problem. 

Trackers go stale because updating them competes with shipping code. When artifacts live in their own frictionless repo, there is nothing competing for attention. A decision or status update takes seconds to commit and lands immediately. 

Delivery repo (no branches, push to main) holds the artifacts; code repos stay feature-branched and PR-gated and point back via CLAUDE.md. 

No connector, no second system 

There is no tracker and no connector. 

The files are the context. The same editor the engineer uses to build is where the assistant reads stories, specifications, decisions, and status. 

That is also why this is a team model rather than a PM productivity hack. 

Managing a project through a Jira connector is something one person does from their own AI session. Here, everyone works from the same source. Engineers build from it. Designers contribute to it. Reports generate from it. 

Nobody waits for a board to catch up. 

The story is a directory, not a ticket 

Our unit of work is a folder, not a line item. 

A story directory holds a few files, each with one job: 

File What it holds 
story.md The contract 
technical-spec.md The plan 
grooming.md The conversation and rationale 
status.md Current state and progress 
decisions/, artifacts/ Evidence and story-scoped decisions 

Together they form the work surface. 

When an engineer opens a story, the assistant sees the same picture they do: the requirements, the implementation plan, prior decisions, open questions, and evidence. 

Dependencies are recorded as simple wikilinks such as [[E2-1]], giving assistants a graph they can traverse rather than a collection of disconnected tickets. 

For engineers: a control plane, not PM docs in Git 

It would be easy to read all of this as project managers moving their paperwork into Git. That is not what it is. 

For an engineer, the repo turns delivery artifacts into part of the work surface, instead of a parallel system to keep in sync. 

Start with the thing engineers feel every day: context switching. Recovering intent from a tracker means leaving the editor, finding the ticket, reading a description written weeks ago, and scanning the comments. Here, story.md, technical-spec.md, and grooming.md sit beside the code. You read the contract, the plan, and the reasoning without leaving Cursor, and so does your agent. 

It also front-loads the ambiguity. The spec-drafting prompt asks the agent to flag acceptance criteria that are vague or untestable before any code gets written, so you are not halfway through a build when you realize the AC could mean two things. 

Evidence lives with the work. artifacts/ is the natural home for eval outputs, screenshots, scorecards, generated datasets, and review forms. When someone revisits the story to debug or review, the proof is right there, not scattered across Drive and Slack. 

Dependencies are plain text. A wikilink like [[E2-3]] is easier to traverse than a tracker’s dependency UI, for you and for the agent, and you can grep the whole tree. 

The code ties back to the contract. Code repos point at the delivery stories, and tests validate the acceptance criteria in story.md, so a reviewer checks against a real definition of done rather than “matches the ticket.” 

The quiet win under all of it: no stale duplicate truth. A tracker usually becomes a second copy of what the engineer learned while building, and it drifts the moment the build teaches you something. Here, when the spec changes, the same repo captures the change, the reason, and the evidence, in one place. 

A lightweight engineering control plane, where context, decisions, specs, evidence, and status live where developers and agents can use them directly. 

The coding agent inherits the contract 

When a coding agent opens a story, it inherits the contract. 

The story, acceptance criteria, technical spec, grooming notes, and code all sit in the same context window. 

Before implementation starts, the engineer has the agent draft technical-spec.md from the story and checks that the plan actually matches the intent. 

Misunderstandings surface in prose instead of three commits later. 

Then the acceptance criteria become something the agent can actively steer toward. 

The definition of done is no longer a memory of a meeting. It is a file sitting beside the code. 

That changes behavior. Agents drift less. They gold-plate less. They spend less time solving adjacent problems and more time solving the one the story actually describes. 

In a tracker-centric workflow, those criteria often live somewhere the coding agent never looks. 

The sprint is a narrative, not a board 

If the story is the unit of work, the sprint is the unit of time. 

One sprint. One file. 

The sprint document contains the goal, committed stories, a running log of mid-sprint events, delivered work, and the retro. 

The key rule is append-mostly. 

Committed work stays committed. When reality changes, and it always does, you add an event explaining what changed and why. 

The sprint becomes a readable history instead of a constantly rewritten snapshot. 

That turns out to be useful for both retrospectives and stakeholder conversations. 

Decisions live where their scope lives 

One small rule removes a surprising amount of friction. 

Story-level decisions live with the story. 

Sprint-level decisions live in the sprint file. 

Project-level decisions live in /decisions/. 

Six months later, when someone asks why a choice was made, there is a dated file with the answer. 

Every ceremony ends in a commit 

The decision lives in the file before the meeting ends. 

If “I’ll update it later” creeps back in, the model degrades into Jira-by-other-means within a sprint. 

Planning, standup, review, and retro all follow the same pattern: 

The ceremony is the conversation. 

The commit is the memory. 

Planning, standup, review, and retro each run as a prompt against the repo and fire a commit before the meeting ends. 

The same data feeds all of these. 

A single status.md file becomes the morning brief, the risk log, the retro, and the stakeholder report. 

You author once and query many ways. 

Status itself remains human-owned. The assistant drafts it from commits and chat context, but a person reviews and commits the result. Otherwise the no-code days (debugging dead ends, credential issues, design discussions) disappear from the record. 

Where it becomes an asset 

Prompts do not stay prompts. 

A query used a few times becomes a skill. 

The skills, plus the directory structure and conventions, become a scaffold. 

The scaffold becomes reusable. 

The next engagement starts from a template that already knows how stories, sprints, decisions, and ceremonies fit together. 

The cost of standing up a new project keeps dropping because the operating layer already exists. 

A prompt used repeatedly becomes a skill, the skills and conventions become a reusable scaffold, and the next engagement starts from that template instead of a blank board. 

One principle keeps that from tipping into over-automation: 

Automate aggregation and synthesis. Do not automate judgment. 

Draft the brief. Draft the retro. Draft the report. 

But decisions, scope, priorities, and commitments stay human. 

The moment you automate judgment, the model quietly rots. 

What it costs, honestly 

It isn’t free. 

The committing discipline is load-bearing, and no tool nags you into doing it. 

It assumes the team works in an agent-aware editor. Open these files in a plain editor and much of the leverage disappears. 

Design tools do not go away. Figma and Miro still exist. The repo simply becomes the place where design decisions and handoffs are recorded. 

There are still open questions too. On a small team without dedicated QA, who formally signs off that acceptance criteria have been met? We are still working that out. 

I’d rather name those edges than pretend the model is finished. 

The honest reframe is that this approach does not eliminate the tracker’s job. 

It absorbs it. 

The consistency Jira enforced for free becomes something you own on purpose. 

For the kind of AI-native work we do, that ownership is the point. 

The repo is already where the work happens. 

So we started asking a simple question: 

What happens if coordination happens there too? 

This operating model is our current answer. 

Not because Git is magical. 

Not because Jira is broken. 

Because operational context now accumulates in one place, and we’re interested in what becomes possible when the people, the artifacts, and the assistants all work from the same source. 

Whether this becomes a broader pattern remains to be seen. 

But after working this way, it is difficult to imagine moving the source of truth back into a system designed primarily to describe the work rather than participate in it.