A3 PORTRAIT · PRINTS IN COLOUR
Source trailReferences

    Runbooks and roadmaps

    How software is run, fixed, planned and delivered, the patterns behind it, and the words that mean something else to engineers.

    Where this volume sits

    Stacks and Sprints runs to two volumes so that each stays short enough to read in one sitting. Every volume holds whole sections in the toolkit’s own order, with the appendices for its own entries. The findings and the explanation of how the toolkit works open volume one. The toolkit closes with the three questions, asked in the team’s own words, at the end of this volume.

    Terms, a drawing and where you’ll meet them

    Part one works differently. Each judging word has its own slide: the definition, the line it usually arrives as in a review, and a drawing of the system as found and as mended. Part three sets each word’s two meanings side by side.

    Words for things

    453 terms and acronyms in 14 families, from the stack to the slang, the patterns teams reuse and the systems large organisations buy.

    Operations and reliability

    What happens after launch: keeping the system up, noticing when it isn’t and learning from failures. These are the words in an incident report or a service contract.

    Each extra nine cuts the downtime allowed tenfold

    Availability is promised as a percentage. The percentage only means something once it is turned back into hours.

    The SLA is promised, the SLO is aimed at, the SLI is measured

    Four terms for one question, how reliable is reliable enough, asked by the contract, the team and the dashboard.

    The same signals serve the team, the person on call and the public

    How a running system is watched: what it records, who looks at it, and who is told when something goes wrong.

    An incident is the event; severity says how bad it is right now

    An incident can begin as a slow page and become an outage within the hour. Severity is graded, and regraded, as it goes.

    The pager wakes one engineer; the commander runs the response

    Who responds to an incident, who leads it, and the written steps they follow.

    A postmortem asks what in the system let it happen, not who

    What follows an incident: a written review, a search for the cause behind the trigger, and a running measure of recovery time.

    Spread the load and keep a spare, so no one part can stop it all

    Four words for surviving load and failure: share the work, grow with demand, keep spares, and find the part that has none.

    RPO is how much data you may lose; RTO is how long you may be down

    When the worst happens, two numbers set the plan: how far back the data goes, and how soon the service returns.

    Plan the work for the quiet weeks and the capacity for the busy ones

    Two kinds of planning around the school year: when downtime is allowed, and how much capacity the peak needs.

    Throttling protects the servers; cutting toil protects the engineers

    Two ways a service is kept going under pressure: slowing the requests on a peak day, and freeing people from the manual work that eats their week.

    Bugs and defects

    The most useful distinction here is severity against priority. Severity is how bad the bug is. Priority is how soon it gets fixed.

    A good bug report says what should have happened and what did

    A bug becomes work when someone writes it down: a ticket with the steps, the expected result and the actual one.

    Severity is how bad the bug is; priority is how soon it is fixed

    Triage gives every new bug two separate ratings, and the two often disagree.

    Not every ticket ends in a fix, and the closing words say why

    Some bugs are documented and lived with; others are closed without a line of code changing.

    The happy path is what the demo shows; edge cases are what schools hit

    Software is built and demonstrated along the expected route. Real use leaves it within a week.

    An error is reported, a crash stops everything, a timeout gives up

    What failure looks like on an engineer’s screen, and the trail it leaves behind.

    A small slip in a loop broke something that used to work

    Two common defects in one picture: a count wrong by one, introduced by a change to code that worked.

    Three faults that freeze, fail or slowly choke a running app

    Three defects that live in the code itself, each with a familiar symptom.

    Bugs that come and go usually depend on timing

    Some bugs appear only sometimes. The cause is often two things happening at once, in an order nobody planned for.

    A hack works today; smells and spaghetti make tomorrow slower

    Three words reviewers use for code that works but will cost more later.

    Software decays even when nobody touches it

    Three ways a codebase ages: code left behind, a world that moves on, and outside code that no longer fits together.

    Planning, rituals and documents

    How teams decide what to build and track it to done. Some of these words, such as velocity and story points, are often misread as measures of productivity when they are only planning aids.

    Waterfall finishes each phase once. Agile repeats them small.

    Three ways to organise delivery, told apart by how often users see working software.

    A sprint opens with a plan and closes with a demo and a retro

    The meetings that give a sprint its shape, from the first morning to the last afternoon.

    An epic is split into stories, and the backlog puts them in order

    Where work waits before it starts, and how a large idea becomes pieces a team can build.

    Acceptance criteria belong to one story. Done applies to them all.

    How a requirement is written, and the checks that decide when it is ready to start and when it counts as finished.

    Sizes come first, points next, and both are forecasts, not promises

    Three ways to say how big a piece of work is, from a rough early guess to a sprint plan.

    Velocity plans the next sprint. KPIs say whether the work mattered.

    Two kinds of number a team reports: one for planning its own work, one for judging results.

    A spike asks one question, and the timebox says when to stop

    How a team buys an answer before it commits to build, without letting the search run on.

    A bug bash hunts for faults; a hackathon builds something new

    Two short, all-hands events: one to break the product before users do, one to try ideas nobody has had time for.

    The board shows where work is, and where it is stuck

    The board most teams work from, and the limits and flags that keep work moving across it.

    Flow measures show where work waits, not how hard people work

    How long work takes to arrive, measured for one card and for a team’s whole delivery.

    A burndown line that climbs means the work grew, not the team slowed

    A sprint chart of work left, and the shape it takes when requirements keep arriving.

    A POC asks if it can be built. An MVP asks if anyone wants it.

    The long view of a product: the plan across quarters, the dated points on it and the early builds that test it.

    A PRD says what to build; a spec, how; an ADR, what was decided

    Four documents that carry a decision from what users need to how it was built and why.

    Who does what

    Who does what in a software team. Titles vary across organisations, so the useful question is less what someone is called than what they decide.

    Two PMs, a PO and a Scrum Master each decide something different

    Four roles that are easy to mix up, told apart by what each one decides.

    A squad holds every skill one product needs, under one tech lead

    Who sits in a typical product team, and who guides the code as against the people.

    Senior engineers can grow in scope without ever managing people

    The senior technical roles, and the career track that lets engineers rise without taking on staff.

    Testers find defects, platform teams build the road, SREs keep it up

    The roles that test the work, build the route to production and keep the live system running.

    Data engineers move the data; analysts and scientists question it

    Three data roles, from the pipelines that move data to the models that run inside an app.

    Research finds the need, and design shapes the flow, look and words

    The roles that decide how a product works, looks and reads, and the role that finds out what users need.

    Leaders set the rules, a BA writes the need, a vendor builds it

    The roles around a software team in a large organisation: leadership, requirements and the suppliers who build.

    Data and AI

    The fastest-moving vocabulary in this glossary. Treat the AI terms as current at October 2026, and check them again before relying on a definition.

    A pipeline carries the data, a warehouse holds it, BI shows it

    Four words for the route data takes from the systems that collect it to the report someone reads.

    Master data is the record everyone copies, so it needs one owner

    Two entries for the core records an organisation runs on, and for the people who answer for them.

    Metadata and lineage let anyone check where a figure came from

    Three entries for describing data well enough that someone else can find it, read it correctly and trace it back.

    Events say what users did. Cohorts and A/B tests make them comparable

    Three words for the usage data software sends home, and for the two ways of making it say something.

    Sign-ups show interest. Activation and retention show real use

    Two entries for following users from first contact to lasting use, and for counting where they drop away.

    One measure to steer by, and a few to check it is not misleading

    Three entries for the numbers at the top of a product’s dashboard: how many use it, what the team steers by, and what users say.

    Accuracy can look high while a model misses the cases that matter

    Two entries for judging a model that sorts things: the answers it is checked against, and three ways of scoring it.

    Training builds the model once. Inference runs it on every request

    Three words for what an AI system is made of, and the two moments in a model’s life that raise different questions.

    A model reads tokens, and only as many as its context window holds

    Four words for what goes into a language model on each request, and the unit every limit and price is counted in.

    Examples show the model the task. Sources let you check its answer

    Three settings in how a model is asked: whether it sees examples, how much it varies, and what it must answer from.

    Fine-tuning changes the model. RAG changes what it reads first

    Three words for making a general model useful on your own material, by two very different routes.

    An agent acts on its own. Guardrails and a person decide how far

    Four words for AI that takes actions, and for the checks around it before and after it acts.

    MCP connects a model to tools. Agentic means it uses them in steps

    Three words for what a current AI system can take in, how it works through a task, and how it reaches other software.

    A copilot suggests code. Vibe coding ships it without reading it

    Two words for writing software with an AI assistant, and for the difference a careful reader makes.

    Open weights let you run a model yourself, on GPUs you pay for

    Two words behind the hosting decision: the chip a model runs on, and whether you may run the model at all.

    A hallucination is wrong once. Model bias is wrong in a pattern

    The two faults of AI output most likely to reach a classroom, and why one is easier to catch than the other.

    A jailbreak is typed by a user. An injection hides in what AI reads

    Two attacks on an AI system, told apart by where the hostile instruction comes from.

    Red-teaming finds the limits a model card should write down

    Three words for testing an AI system before students use it, and for the document that records what testing found.

    Shorthand and slang

    The words of chat channels and review comments. Several of them carry more than they seem to: LGTM is a sign-off, “nit” means the comment is optional, and bikeshedding is a criticism of the meeting you’re in.

    LGTM approves. A nit is optional. WIP means do not review yet

    Five words that tell everyone where a code review stands, from not ready to approved.

    Thread shorthand says how sure someone is, and whether they agree

    Five pieces of shorthand from proposal threads, each saying something about the writer’s certainty or support.

    Three ways to ask for someone’s time, and two polite ways to say no

    Words for asking for attention and for protecting it, as they appear in a team channel.

    Greenfield starts clean. Brownfield lives with what is already there

    Words for where a new system is built, and for the old one it is meant to replace.

    Know how far a mistake can reach, and which route avoids it

    Three words for the risk in a change: how far it spreads, the safe road, and the trap beside it.

    A system is safer when more people know it and its makers use it

    Three words for how knowledge of a system is spread, tested and handed on.

    Yak shaving and bikeshedding both spend the day on the wrong thing

    Two words for time that disappears: into a chain of side tasks, or into the easiest decision in the room.

    YAGNI, DRY and KISS all argue for less, in different ways

    Three acronyms heard in design reviews, each a short argument for a smaller design.

    Four things said in bug threads, and the one that finds the bug

    Four phrases from bug threads, from the unhelpful to the one that actually finds the cause.

    Patterns and strategies

    Named, reusable answers to problems every software team meets. Most share one idea: shrink the bet.

    Modernising legacy systems

    Ten terms on four plates, from strangler fig pattern to cutover.

    Retire an old system one route at a time, not all on one date

    Two ways to replace an old system: move everything on one day, or move one function at a time until nothing is left behind.

    Six verdicts for each old system. Lift and shift is only one.

    A migration is a sorting exercise: each legacy system gets one of six verdicts, and moving it unchanged is only the first.

    Put a steady front in place, then change what sits behind it

    Three ways to put a fixed surface between the people who use a system and the parts being changed behind it.

    Load the past, run both, compare, and only then switch

    Three steps that turn a migration from a leap into a checked handover: history first, both systems side by side, then one switch.

    Slicing and sequencing delivery

    The source glossary groups this family’s terms under five headings, and the plates follow them. This group’s plates:

    A thin slice works early. A layer works only when all are done.

    The same work cut two ways: down through every layer for one feature, or across one layer for every feature.

    Three ways to get something working end to end before it is good

    Three close cousins: all build end to end first, and they differ in what they connect and what they keep.

    Two ways to grow the work, and a roadmap that admits uncertainty

    How work grows, in rounds over the whole or in finished pieces, and how to show a plan without promising dates nobody can keep.

    Releasing safely

    Eleven terms on four plates, from dark launch to chaos engineering.

    Try it on the live system where nobody sees it, with a way out

    Three ways to put new work on the live system while keeping users safe from it until it has proved itself.

    Find the problem while it is small, early and cheap to fix

    Three habits that move discovery earlier: small changes merged often, checks done sooner, and failures surfaced at once.

    Build up from a base that works, or plan the fall back to one

    Two routes to the same promise: the core task still works when the network, the device or a service does not.

    Assume parts will fail, and design so that a failure stays small

    Three defences against failure that spreads: repeats that are safe, a breaker that isolates, and drills that prove both.

    Architecture and team design

    Fourteen terms on six plates, from conway’s law to infrastructure as code (IaC).

    The system will copy the org chart unless the teams are designed

    A system’s seams fall where its makers’ seams are, so the shape of the teams is a design decision in its own right.

    Most teams should own a journey; the rest exist to serve them

    Four kinds of team, each with one job, all kept small enough that everyone in them can still talk to everyone else.

    One word, one meaning, inside a boundary everyone can see

    Software built in the language of the people who do the work, with a clear line where one meaning of a word stops.

    Keep each fact in one place, and announce it once when it changes

    Three ways to stop data drifting apart: one home for each fact, parts with separate jobs, and changes broadcast once.

    Write it once, and choose which hard case to design for first

    Content kept apart from any one screen, and a choice about which hard case the design must meet before any other.

    Know what you will own, and write down how it was set up

    Two decisions about ownership: whether to make the software at all, and whether its set-up can be rebuilt from a file.

    Anti-patterns

    Eleven terms on four plates, from big ball of mud to brooks’s law.

    The mess invites a rewrite, and the rewrite tries to fix everything

    Two failures that often arrive as a pair: a system with no shape left, and a replacement asked to carry every wish at once.

    Answers chosen by habit, pride or imitation

    Three ways a team picks a solution for reasons that have nothing to do with the problem in front of it.

    Effort that looks like progress and changes nothing for users

    Three ways a busy team spends its time on things nobody needed: polish, speed and a count of features.

    Three ways to run out of time: no decision, no scope, more people

    A decision studied forever, a date fixed before the work is known, and staff added to a project that is already late.

    Buying and running enterprise software

    The words of procurement and of the systems large organisations run on.

    Ask the market, compete the bid, then live by the statement of work

    The documents of buying software in the order they appear: asking the market, inviting bids, signing the scope and changing it.

    The licence is the price you see; the rest of the bill comes later

    What ready-made software really costs: the licence, the shape of its bill over time, and everything that surrounds it.

    The systems IT knows about, and the tools staff found for themselves

    The large systems an organisation runs on, and the tools that grow up beside them when those systems do not meet a need.

    One front door for help, and a clear route up when it cannot fix it

    How help is organised: a service desk that takes every request, and tiers that pass the hard ones to people who can fix them.

    Where non-engineers get lost

    Words that mean one thing in education and policy and another in engineering, and the long road from “merged” to “done”.

    Same word, different meaning

    These words are used every day in education and policy, and they mean something else to engineers. Meetings go wrong when both sides nod at the same word and hear different things.

    Words for structures

    Five words for how things are organised. In a ministry they name ideas and programmes. In engineering they name code and copies of it.

    Words for moving things

    Five words for moving something to where it is used. Every one of them names a different kind of movement in a school.

    Words for the work itself

    Five words for planning and finishing work. They are the ones most likely to appear in a status update read by both sides.

    Words for checking

    Three words for judging whether something is good enough. Education checks people. Engineering checks code and specifications.

    Words for people and ties

    Four words for who someone is and what something rests on. Two of them flip from praise to warning.

    From merged to done

    “Done” is the riskiest word in a status update. Between a change being merged and everyone having it, there are as many as seven distinct steps, and each has its own word.

    Users see nothing until step six

    Automated checks pass. Nobody outside the team can see it.

    A feature flag can keep finished work dark for weeks

    Deployed and released are two events, not one. The flag between them is what makes a release reversible, and it is also why “it’s in prod” does not mean anyone can use it.

    When someone says it’s done, ask which step.

    Merged, built and deployed are engineering milestones. Released and generally available are the ones users notice. Teams vary: some skip UAT, and some release the moment they deploy. The question still works.

    Ask the three questions in the words the team already uses

    None of these questions needs technical knowledge to ask, and each of them has a factual answer. That is the test of a good question to an engineering team: it can be answered by looking something up.

    Thirteen works

    The definitions are the source glossary’s own. Where a term rests on a published source, the source is named in the definition and listed here.

    Every term in this volume is indexed, with its definition, at https://www.geraldajam.com/library/glossary/phrasebooks/stacks-and-sprints.