In his essay on Salvador Dali,
Orwell argued that because Dali was a repulsive human being,
the right wouldn’t admit that he was a great artist;
conversely, because he was a great artist,
the left wouldn’t admit he was a repulsive human being.
I’m seeing the same thing with AI:
because it’s unethical,
one side won’t acknowledge that it’s useful,
but because it’s useful,
the other side won’t acknowledge that it’s unethical.
Thirty years ago,
when I first started writing for Doctor Dobb’s Journal,
I decided to do a piece on the object-oriented features that were being added to MATLAB 5.
I didn’t know much about them, so I called The MathWorks,
told the publicity rep who I was,
and asked if they could find me a co-author.
A couple of days later I got a call back from a woman who said she had volunteered to help with the article.
I assumed she was also from marketing,
so I explained again who I was and what I wanted.
When she said she’d be happy to provide me with information and examples,
I replied,
“Thanks, but I’d rather have someone technical.”
After a slight pause, she said,
“Well, I have a master’s degree in Computer Science, and I implemented some of the new features.”
I apologized,
but the next day
I had an email from someone else saying that he’d be working with me.
Time for another cup of tea.
If you came in peace,
be welcome.
NOAA’s Dealing with Disruptive Behaviors
describes ten different kinds of people who can disrupt meetings and gives advice for handling each.
I’ve been a big fan of this taxonomy since I first encountered it several years ago,
but going through it again yesterday,
I realized that there’s a fair bit of overlap between some of the types.
In order to slim it down,
I grouped them as follows:
Aggressive
Assertive
Passive
People-focused
Blowfish, Dolphin
Sea Otter, Flounder
Task-focused
Jellyfish, Shark
Sea Lion
Crab, Clam, Octopus
I then picked one from each group,
looked at the result,
and added Clam back into the mix
because I felt this type was distinct enough from Crab to merit inclusion.
This came up as part of an exercise to analyze
how people are responding to AI mandates (both positively and negatively)
and how to manage those responses.
If you have seen other work that looks at this,
or have feedback on how I’ve slimmed these categories down,
I’d be grateful for feedback.
Talkative Blowfish
Blowfish is a very chatty, assertive people person.
Blowfish wants everyone to feel comfortable and positive about the process.
They have a tendency to be overly talkative (almost compulsively)
because they are enthusiastic, want to show off, or are well-informed and eager to use their knowledge.
Blowfish can dominate the floor time at the expense of other group members.
While they frequently have good ideas and strong contributions to make,
they also ramble, monopolize the discussion, and do not give others an opportunity to express their thoughts.
Subsumes Dolphin: Blowfish rambles about the work, Dolphin diverts from it.
Eager Sea Otter
The sea otter is a passive people person and wants everyone to get along.
Sea otters are super agreeable, overly positive people
who are optimistic, very reasonable, sincere, and supportive.
They are people-oriented and aim to please those nearby (for instance, by always saying yes).
They seek approval by giving approval.
This may cause difficulty in group situations because they overcommit or are unreliable.
Subsumes Flounder: the Sea Otter is eager to please, not detached.
Dominating Shark
The shark is aggressive and focused on efficiency and the task.
They can be hostile and dominating and might try to intimidate and bully people.
They make cutting remarks or throw temper tantrums when they do not get their own way.
Some hostile individuals will be task-focused and want to get the job done while maintaining control.
These individuals will generally have a more focused attack
on the failure of others to complete a specific task or take necessary actions.
Others may explode and attack other people in a more random fashion,
which is typically done to command attention.
Subsumes Jellyfish: the Shark’s aggression is about domination, not intellectual sport.
Complaining Crab
The complainer can come in many forms: whiner, critic, or obstructionist.
Crabs are passive and task-focused, and they want to get it done.
Despite the negative connotation, this person is often motivated by perfection.
Negative, complaining people may seem to object to everything,
asserting that ideas proposed will not work or are impossible.
The complainer may completely deflate any optimism others express for a project
and may block others from accomplishing goals.
Crabs gripe and do little to improve the situation,
either because they feel powerless
or because they refuse to bear the responsibility for an imperfect solution.
Subsumes Octopus: the Crab complains outward rather than freezing inward.
Arrogant Sea Lion
Sea lions are assertive and need the group to accept their expertise,
and they can become know-it-alls when questioned.
They believe that they have more credibility than has been acknowledged
and want everyone to understand and agree with them.
The sea lion knows a lot about the topic but does not contribute in a way that sits well with other participants,
sometimes using their credentials, age, length of service, or residency to disparage an idea.
With cockiness and an inflated ego,
the sea lion can be condescending, imposing, pompous, or arrogant toward others.
In all likelihood, this behavior will make others feel as though there is no point in contributing.
Shy Clam
The clam is shy and quiet, passive, and task-focused.
Clam wants to get it right.
Shy individuals may be reluctant or afraid to express their ideas in a group setting,
so they may appear to be unresponsive.
I’m co-teaching a lesson for the Carpentries next week
about the impact of LLMs on teaching.
Here are a few things I’ve been reading to prepare:
Barba2026
Lorena A. Barba and Laura Stegner:
“The Conversational Exam: A Scalable Assessment Design for the AI Era”.
https://arxiv.org/abs/2601.10691,
2026.
Conversational exam (live coding + explanation in small groups) restores assessment validity against generative AI cheating; 58 students examined in 2 days; combines authentic practice with inherent validity.
Bielaczyc1995
Katerine Bielaczyc, Peter L. Pirolli, and Ann L. Brown:
“Training in Self-Explanation and Self-Regulation Strategies: Investigating the Effects of Knowledge Acquisition Activities on Problem Solving.”
Cognition and Instruction.
13(6),
1995.
https://doi.org/10.1207/s1532690xci1302_3.
Training study (24 novice programmers) showing self-explanation and self-regulation strategy training causally improves programming task performance; instructional group showed significantly greater strategy use and performance gains.
Bridgeford2025
Eric W. Bridgeford, Iain Campbell, Zijao Chen, et al.:
Ten Simple Rules for AI-Assisted Coding in Science.
https://arxiv.org/abs/2510.22254,
2025.
10 practical rules for AI-assisted coding in scientific computing; addresses problem preparation, context management, testing/validation, and code quality; emphasizes human agency and domain expertise for reproducible research.
Butler2024
Jenna Butler, Jina Suh, Sankeerti Haniyur, and Constance Hadley:
“Dear Diary: A Randomized Controlled Trial of Generative AI Coding Tools in the Workplace.”
https://doi.org/10.48550/arxiv.2410.18334,
2024.
Mixed-methods study (survey + RCT + 3-week diary) on generative AI coding tools at a large multinational; sustained use increases perceived usefulness and enjoyment; trustworthiness perceptions unchanged; 84% report positive daily work changes; unexpected uses include web search replacement.
Deslauriers2019
Louis Deslauriers, Logan S. McCarty, Kelly Miller, Kristina Callaghan, and Greg Kestin:
“Measuring Actual Learning Versus Feeling of Learning in Response to Being Actively Engaged in the Classroom.”
Proc. National Academy of Sciences,
116,
Sept. 2019.
https://doi.org/10.1073/pnas.1821936116.
RCT shows active learning produces more learning but lower perceived learning; increased cognitive effort is misread as poorer learning; early instructor intervention corrects this misperception.
FinnieAnsley2022
James Finnie-Ansley, Paul Denny, Brett A. Becker, Andrew Luxton-Reilly, and James Prather:
“The Robots Are Coming: Exploring the Implications of OpenAI Codex on Introductory Programming.”
Proc. 24th Australasian Computing Education Conference,
https://doi.org/10.1145/3511861.3511863,
2022.
OpenAI Codex outscores most students on intro programming exams; handles Rainfall problem variants well; generates diverse solutions for identical prompts; raises challenges and opportunities for CS education.
Jiao2026
Yuling Jiao and Qiuli Wang:
“Large language models for formative feedback in writing instruction: a systematic review of classroom interventions, feedback quality, and student outcomes”.
Frontiers in Education,
11,
2026,
https://doi.org/10.3389/feduc.2026.1834085.
Studies in which teachers discussed AI-generated feedback, helped students interpret it, or combined it with their own comments generally reported better learning outcomes than studies where students worked independently with AI.
Leinonen2023a
Juho Leinonen, Paul Denny, Stephen MacNeil, et al.:
“Comparing Code Explanations Created by Students and Large Language Models.”
Proc. 2023 Conference on Innovation and Technology in Computer Science Education,
https://doi.org/10.1145/3587102.3588785,
2023.
LLM-generated code explanations are rated significantly more accurate and understandable than student-generated ones in a 1000-student course; scalable on-demand explanations can scaffold introductory programming learning.
Leinonen2023b
Juho Leinonen, Arto Hellas, Sami Sarsa, et al.:
“Using Large Language Models to Enhance Programming Error Messages.”
Proc. 54th ACM Technical Symposium on Computer Science Education,
https://doi.org/10.1145/3545945.3569770,
2023.
LLMs enhance Python error messages with plain-language explanations and fix suggestions; sometimes surpass original messages in interpretability and actionability for novice programmers.
Ma2024
Qianou Ma, Hua Shen, Kenneth Koedinger, and Sherry Tongshuang Wu:
“How to Teach Programming in the AI Era? Using LLMs as a Teachable Agent for Debugging.”
Lecture Notes in Computer Science,
https://doi.org/10.1007/978-3-031-64302-6_19,
2024.
HypoCompass trains students to debug LLM code by having them hypothesize error causes while LLMs handle code completion; improves debugging performance 12% over pre-test with fourfold efficiency vs. human tutors.
Ma2025
Qianou Ma, Weirui Peng, Chenyang Yang, Hua Shen, Ken Koedinger, and Tongshuang Wu:
“What Should We Engineer in Prompts? Training Humans in Requirement-Driven LLM Use.”
ACM Transactions on Computer-Human Interaction,
32(4),
https://doi.org/10.1145/3731756,
2025.
Randomized experiment with 30 novices finds Requirement-Oriented Prompt Engineering (ROPE) training achieves 20% gains vs. 1% for conventional prompt engineering training.
OBrien2026
Gabrielle O’Brien, Alexis Parker, Nasir Eisty, and Jeffrey Carver:
“A survey of generative AI adoption and perceived productivity among scientists who program.”
2026,
https://doi.org/10.48550/arXiv.2512.19644.
Survey of 868 scientists who program as part of their work,
reporting that 80% use GenAI tools in their programming,
with 77.5% of those using general purposing tools like ChatGPT over specialised coding tools.
Peng2023
Sida Peng, Eirini Kalliamvakou, Peter Cihon, and Mert Demirer:
“The Impact of AI on Developer Productivity: Evidence from GitHub Copilot.”
2023,
https://doi.org/10.48550/arXiv.2302.06590.
Randomized controlled experiment claiming that GitHub Copilot
users completed a JavaScript coding task 55.8% faster than the
control group.
Richards2026
Jonan Richards, Bruno Alves de Oliveira, Iury Oliveira, Igor Wiese, and Mairieli Wessel:
“No Two Developers Think Alike: How Problem-Solving Styles and Experience Shape Needs in Conversational Interaction with Copilot”.
2026,
https://arxiv.org/abs/2606.19216.
Characterizes 5 distinct interaction modes and 10 underlying needs in developers’ interactions with AI tools.
Sadowski and Zimmerman 2019
Caitlin Sadowski and Thomas Zimmermann (eds.):
Rethinking Productivity in Software Engineering.
Apress,
2019,
9781484242216.
Edited volume collecting research and practitioner perspectives on
how to understand, define, and measure software developer
productivity.
Stray2026
Viktoria Stray, Elias Goldmann Brandtzæg, Viggo Wivestad, Astri Barbala, and Nils Brede Moe:
“Developer Productivity with and Without GitHub Copilot: A Longitudinal Mixed-Methods Case Study.”
Proceedings of the 59th Hawaii International Conference on System Sciences,
https://doi.org/10.24251/hicss.2026.880.
2026.
Mixed-methods study of 703 NAV IT repositories finds Copilot users were more active even before adoption and shows no statistically significant changes in commit-based metrics after adopting the tool.
I finally have time to flesh out ideas for lessons that I’ve wanted for years.
However,
if I can’t find a way to send them back to 2006,
there’s no point writing them:
very few people read long-form tutorials about software these days.
I’d still be interested in comments, though—figuring out what I would teach
always helps me learn.
Overview
Topic: Error handling.
Audience: Senior undergraduates who are
comfortable writing programs in Python and JavaScript that are hundred of lines long,
know how to raise and catch exceptions,
and are familiar with SQL, C, and the Unix shell,
but have no experience building production-robust programs.
Format: seven 45-minute lessons with exercises.
1) What Can Go Wrong
Taxonomy of errors
Logic errors: bugs in the program itself
Runtime errors: null dereference, division by zero, index out of bounds
External errors: file not found, network timeout, database constraint violation
Human errors: bad input, misconfiguration, wrong file format
Environmental errors: disk full, out of memory, clock skew
Failure modes
Fail-fast vs. fail-slow: silent corruption is harder to debug than an early crash
Silent failures: errors ignored, wrong results returned without warning
Cascading failures: one component’s error triggering failures in others
C-style error signaling
Return codes: functions return -1, NULL, or 0 on failure
errno: a global (thread-local) integer set by system calls
Read with perror() or strerror()
Sentinel values: EOF (-1), invalid index, or a special out-of-band value
Advantages: explicit control flow, no hidden jumps, zero runtime overhead
Disadvantages: easy to ignore, verbose, callers must check every call, no stack information
Common pitfalls:
Not checking return values
Reading errno after another call has overwritten it
errno not being set on success
Philosophy: errors are not exceptional, they are expected
The happy path is one of many paths
Spectrum of responses: ignore vs. crash vs. recover vs. degrade gracefully
Choosing a response requires knowing the context and the cost of each option
Taxonomy drill
Given six short programs in Python, JavaScript, and C,
each containing a different type of error,
classify each error using the taxonomy above
and explain whether the program’s response (crash, wrong output, hang, silent skip)
is appropriate for a production context.
errno pitfall hunt
A short C function uses fopen, fread, and fclose and checks errno after each call.
The function contains three bugs related to C-style error handling
(e.g., ignoring a return value, checking errno too late, not distinguishing error from end-of-file).
Identify each bug and propose a fix.
Failure brainstorm
Given a brief description of a web form that accepts a user’s name, email, and a file upload,
then stores the data in a database,
list every error that could occur,
including errors the user causes,
errors the network causes,
errors the OS causes,
and errors the program itself could cause.
Compare lists with a partner and identify any category you missed.
2) Error Propagation and Recovery
Options when an error is detected
Crash/abort: call abort(), panic, or let the process die
Appropriate when the program is in an unrecoverable state or an invariant is violated
Raise an exception: hands control to the caller’s handler
Log and continue: almost always wrong
Hides failures and corrupts program state
Retry: attempt the operation again (only safe for transient, idempotent operations)
Use a fallback value: return a default, a cached result, or a degraded response
Partial success: complete what you can, report what failed
Compensating action: undo work already done before propagating the error
Propagating errors without losing information
Re-raise: pass the error up unchanged
Wrap/chain: add context while preserving the original cause (Python raise X from Y, Java initCause)
Translate: convert a low-level error into a domain-level error
(e.g., from OSError to ConfigurationError)
Anti-pattern: catching an exception and throwing a new one that discards the original
Adding context as errors propagate
Include what was being attempted and with what inputs (sanitized)
Each layer adds the context it has
Callers should not have to guess
C-style propagation patterns
Check every return value, every time: a missed check is a latent bug
Passing error information up via out-parameters or a shared error struct
goto cleanup pattern for resource cleanup on multiple error paths
Comparing C-style propagation to exception propagation: explicitness vs. verbosity
Retry logic
Distinguish transient errors (timeout, temporary unavailability)
from permanent errors (permission denied, not found)
Exponential backoff: double the wait time between retries
Add jitter (random variation) to prevent thundering herd
Set a maximum number of retries and a total timeout
Only retry idempotent operations
Retrying a non-idempotent operation can cause duplicate side effects
Propagation rewrite
A Python function reads a configuration file,
connects to a database using values from that file,
and runs a query.
The function is given with no error handling.
Rewrite it so that errors at each stage are wrapped with context and propagated appropriately,
then rewrite the equivalent in C using return codes, noting what is harder and easier in each style.
Swallowed exceptions
A code snippet contains three except: pass or equivalent constructs.
For each, explain what failure is being hidden,
what could go wrong as a result,
and what the correct handling would be.
At least one case should be a place where logging-and-continuing is wrong even though it feels safe.
Retry with backoff
Implement a retry decorator (Python) or higher-order function (JavaScript) that wraps a function call,
retries up to N times on specified exception types,
uses exponential backoff with jitter,
and raises the last exception if all retries fail.
Test it against a stub that fails a configurable number of times before succeeding.
Discuss which HTTP status codes should trigger a retry and which should not.
3) Error Communication
Errors have multiple audiences with different needs
End users: need to know what happened and what they can do about it
API callers: need structured, machine-readable information to handle programmatically
Operators and on-call engineers: need enough detail to diagnose and fix the problem
User-facing error messages
State what went wrong in plain language
Tell the user what to do next (retry, contact support, correct their input)
Avoid technical jargon, stack traces, and internal identifiers
Do not blame the user
Provide a reference (request ID, error code) they can give to support without revealing internals
Distinguish “you did something wrong” (4xx) from “we did something wrong” (5xx)
API error responses
HTTP status codes: 4xx for client errors, 5xx for server errors
Use specific codes (400, 401, 403, 404, 409, 422, 429, 503) rather than always returning 400 or 500
Error response body: include type, message, detail, and request_id fields
RFC 7807 (Problem Details for HTTP APIs): a standard format worth knowing
Consistent schema across all endpoints
Callers should not have to handle different shapes
Validation errors: return all field errors at once, not just the first one
Security considerations in error communication
Information leakage:
stack traces, SQL queries, file paths, and internal service names in error messages help attackers
User enumeration: “user not found” vs. “wrong password” reveals whether an account exists
Use “invalid credentials” instead
Timing side-channels: if “user not found” returns in 1ms and “wrong password” returns in 200ms
(due to password hashing),
an attacker can enumerate accounts by timing
Always perform the same operations regardless of the error branch
Error messages as reconnaissance:
the more specific the error, the more an attacker learns about your system
Rule of thumb:
the user-facing message and the internal log entry should contain different levels of detail
Correlation IDs
Generate a unique ID per request
Include it in all logs and in the user-visible error response
Allows an operator to find all logs related to a specific error report
Pass the ID through every internal service call in a distributed system
Message rewrite
Five error messages from real or realistic applications are provided
(e.g., a raw Python traceback shown in a web UI,
a message reading “MySQL error 1045 for user ‘admin’@’localhost’”,
a message reading “Your session token has expired and cannot be refreshed”).
For each,
write a user-appropriate replacement that is informative without leaking internals,
and write a separate message suitable for the internal log.
API error schema design
Design the error response body for a REST endpoint that registers a new user.
The schema must handle missing required fields,
fields that fail validation (email format, password length),
a duplicate username,
and an unexpected server failure.
Show example JSON for each case.
Login timing attack
A login function is provided that queries the database first
and returns “User not found” immediately if the user does not exist,
or hashes the password and compares it (taking ~200ms) if the user does exist.
Explain the timing side-channel,
fix it so both branches take the same time,
and discuss whether this fix is always necessary or whether it depends on the threat model.
4) Logging and Observability
Why logging matters for error handling
Errors that are caught and handled still need to be visible to operators
Post-mortem debugging requires a record of what happened
Trends in error rates reveal systemic problems before they become outages
Log levels and when to use each
DEBUG: detailed internal state, useful during development, not in production by default
INFO: normal significant events (startup, shutdown, completed request)
WARNING: unexpected but handled conditions that may indicate a problem
ERROR: a failure that was caught but means something did not complete successfully
CRITICAL: system is in a severely degraded or unrecoverable state
Common mistake: logging every caught exception at ERROR regardless of severity
Structured logging
Free-text logs are hard to query: key-value pairs (or JSON) are machine-parseable
Standard fields: timestamp, level, message, request_id, user_id, error_type, duration_ms
Log the event, not a sentence:
{"event": "db_query_failed", "table": "users", "error": "timeout"}
rather than "Failed to query users table due to timeout"
What to include in an error log entry
What was being attempted and with what parameters (sanitized)
The error: type, message, and stack trace for unexpected errors (not for expected/handled errors)
The outcome: did the system recover, degrade, or fail?
Correlation ID to link this entry to the request and to other services
What NOT to log
Passwords, API keys, session tokens, OAuth codes: not even partial values or hashes
Personally identifiable information (PII): names, emails, government IDs
Check compliance requirements (GDPR, HIPAA)
Full request/response bodies if they may contain credentials or sensitive data
Stack traces in user-facing API responses (they belong in the internal log only)
Audit logs vs. diagnostic logs
Diagnostic logs: help engineers debug problems (mutable, short retention, lower protection)
Audit logs: immutable record of security-relevant actions
(who authenticated, what data was accessed, what was changed)
Long retention, tamper-evident, access-controlled
Not the same stream: conflating them causes both compliance and debugging problems
Traces: end-to-end path of a request through multiple services
Correlation IDs connect all three (errors visible in metrics can be traced to specific log entries)
Logging retrofit
A function that processes uploaded CSV files
is given with a single except Exception as e: print(e) handler.
Add structured logging using Python’s logging module (or a JavaScript equivalent).
Include appropriate log levels for different error conditions,
add relevant context fields,
distinguish expected errors (bad CSV format) from unexpected errors (disk full),
and identify what was unobservable before your changes.
Security audit of log output
Three log excerpts are provided,
each containing at least one security violation
(e.g., a plaintext password,
a full stack trace that reveals an internal file path and SQL query,
or a user ID that allows enumeration of account existence).
Identify each violation, explain what risk it creates, and show the corrected log entry.
Logging strategy design
A data pipeline reads records from an API, transforms them, and writes results to a database.
The pipeline processes records in batches.
Design a logging strategy: what events are logged at each stage,
what fields each entry includes,
what constitutes a loggable error vs. a metric,
and how an operator would diagnose a job that completed without crashing
but produced fewer output records than expected.
5) External Systems
Why external systems are the primary source of production errors
Code you write can be tested exhaustively
Systems you depend on cannot be controlled
External systems fail in ways that are partial, delayed, and inconsistent
File system errors
Not found, permission denied, disk full, file locked by another process, corrupted data
Always close files: use context managers (with in Python, using in C#, RAII in C++)
Atomic writes: write to a temporary file, then rename
Rename is atomic on POSIX systems
Direct writes leave a window where the file is partially written
Handling partial reads and partial writes explicitly
Database errors
Connection failure: the database is unreachable (transient, so retry with backoff)
Timeout: query took too long (may or may not have committed, so check before retrying)
Constraint violation: duplicate key, foreign key failure, not-null violation (permanent, so do not retry)
Deadlock: two transactions are waiting on each other (transient, so retry the whole transaction)
Serialization failure (in serializable isolation):
transaction conflicted with a concurrent one (transient, so retry)
Connection pool exhaustion: all connections are in use (application-level backpressure needed)
Distinguishing transient from permanent errors using error codes, not message strings
Network and HTTP errors
Every external call must have a timeout: separate connect timeout and read/total timeout
Retrying safely:
only retry if the operation is idempotent or the error occurred before the request was received
HTTP status codes that indicate retrying is safe: 429 (with Retry-After), 503, 504
HTTP status codes that should not be retried: 400, 401, 403, 404, 409, 422
Exponential backoff with jitter (review from Lesson 2, now applied to real HTTP clients)
Never assume the schema of data from an external source
Validate presence and type of required fields before using them
Handle missing optional fields explicitly rather than letting a KeyError or undefined propagate
Version skew: the API changed its schema (your parser must detect and report this)
Fail loudly on unexpected schema changes rather than silently producing wrong results
Error handling retrofit for a pipeline
A script that downloads a JSON file from an HTTP endpoint,
parses it,
and inserts records into a SQLite database is given.
It has no error handling.
Add appropriate handling for
HTTP errors (including distinguishing retryable from non-retryable),
JSON parse failures,
missing required fields,
and database constraint violations.
For each error type, decide whether to abort, skip the record, or retry, and justify the choice.
Transaction retry
A function executes a multi-statement database transaction and fails with a deadlock error.
Implement retry logic that retries the entire transaction on deadlock,
does not retry on constraint violations,
limits total retries,
and logs each retry attempt.
Discuss why retrying only the failed statement, rather than the whole transaction, is wrong.
Atomic file write
A function saves user settings to a JSON file by opening the file,
serializing the settings,
and writing directly.
Demonstrate two failure scenarios where this approach corrupts the file:
interrupted write,
and crash between open and close.
Implement the temp-file-then-rename pattern and explain why it is safe
even if the process is killed mid-write.
6) Concurrency and Async Errors
Why concurrency changes error handling
An error in a thread or task may not propagate to the code that started it
Operations may be partially complete when an error occurs, leaving shared state inconsistent
Some errors (race conditions, deadlocks) only appear under concurrency and are hard to reproduce
Errors in threads
Python: an uncaught exception in a Thread prints a traceback
but does not crash the main thread or propagate to the caller
Errors are silently lost unless explicitly collected
Thread pools (concurrent.futures.ThreadPoolExecutor):
exceptions are stored and re-raised when future.result() is called
But only if you call it
Always collect results from thread pools
Never fire-and-forget threads that could fail silently
Setting a global thread exception handler (threading.excepthook) as a backstop
Async/await errors (Python and JavaScript)
Python: an unawaited coroutine silently does nothing (forgetting await is a silent bug)
Python: an unhandled exception in a Task is reported when the task is garbage-collected
Too late to be useful
Always await tasks or attach .add_done_callback
JavaScript: an unhandled promise rejection crashes Node.js in recent versions
In older versions (and some browsers) it is silently swallowed
Always handle rejections: .catch(), try/await/catch, or process.on('unhandledRejection')
Forgetting await in a loop:
tasks are created but not awaited, the loop exits, and all tasks are cancelled
Parallel operations and partial failure
Promise.all: fails fast
If any promise rejects, the whole thing rejects and other results are lost
Promise.allSettled: waits for all promises regardless of outcome
Returns an array of {status, value/reason}
Use when you want partial results
Python asyncio.gather(return_exceptions=True):
returns exceptions as values instead of raising them
Allows processing partial results
Decision: whether to fail fast or collect all results depends on whether partial results are useful
Structured concurrency
The problem: a task that spawns subtasks and then throws has left orphaned subtasks running
Python asyncio.TaskGroup (3.11+):
all tasks in the group are cancelled if any raises an exception
The group does not exit until all tasks are done
Structured concurrency ensures the lifetime of every subtask is bounded by the scope that created it
Cancellation propagation:
when a task is cancelled, it receives a CancelledError (Python) or AbortError (JS)
Cleanup must happen in finally or try/catch around await
Shared state and locks in error paths
A lock must always be released, even if an error occurs
Use try/finally or a context manager
If an error occurs while holding a lock after modifying shared state,
the state may be inconsistent when the lock is released
Decide if the modification needs to be rolled back before releasing
Deadlock: two threads each hold a lock the other needs
Prevent via lock ordering (always acquire locks in the same order) or via timeouts
Silent thread failures
A Python script spawns ten threads,
each of which downloads a file and writes it to disk.
When a download fails,
the exception is printed to stderr but the main thread sees all tasks as complete.
Rewrite using ThreadPoolExecutor,
collect all futures,
and report which downloads succeeded and which failed,
then introduce a deliberate failure in two of the threads
and verify the errors are caught and reported correctly.
Promise.all to Promise.allSettled
A JavaScript function fetches data from five independent APIs using Promise.all.
The whole function fails if any one API is unavailable.
Rewrite it using Promise.allSettled so that
results from available APIs are returned and failures are reported per-API.
When is Promise.all’s fail-fast behavior preferable?
Deadlock identification and fix
A function managing a shared cache acquires cache_lock and then stats_lock in that order.
Another function acquires the same locks in the opposite order.
Trace through a scenario where both functions run concurrently and deadlock.
Fix the bug using lock ordering.
As a second fix,
add a timeout to the lock acquisition and show how to handle the case where the timeout expires.
7) Resilience Patterns and Production Practices
The goal: systems that degrade gracefully rather than fail catastrophically
“Fail safe” vs. “fail secure” vs. “fail operational”
No system achieves zero errors: the goal is bounded, predictable failure
Circuit breaker pattern
Closed state: requests pass through normally
Open state: requests fail immediately without attempting the operation
(protecting a downstream that is already failing)
Half-open state: a probe request is allowed through
If it succeeds, the breaker closes
If not, it stays open
Thresholds: open after N failures in a time window
Reset after a timeout
Use case:
preventing cascading failures when a slow or failing dependency causes your thread pool to fill up
Bulkhead pattern
Isolate resources (thread pools, connection pools, processes)
so that a failure in one area does not exhaust resources for another
Example: use separate HTTP connection pools for critical and non-critical external services
Trade-off: resource isolation requires over-provisioning in aggregate
Timeout patterns
Every call to an external system must have a timeout
A call without a timeout can block forever
Cascading timeouts:
internal timeouts should be shorter than the external deadline so the caller still gets a response
Distinguishing timeout-on-connect from timeout-on-read
Graceful degradation
Serve stale cached data rather than failing when the data source is unavailable
Disable non-critical features (recommendations, analytics) when load is high or dependencies are down
Return partial results rather than failing when some data is unavailable
Read-only mode: allow reads but reject writes when the system cannot safely persist data
Testing for failure
Happy-path tests verify normal behavior
Error-path tests verify that errors are handled correctly
Chaos engineering: inject controlled failures in production
to find weaknesses before an actual incident does
Error budgets and SLOs
Service Level Indicator (SLI): a measurement of reliability (e.g., fraction of requests that succeed)
Service Level Objective (SLO): the target (e.g., 99.9% success over 30 days)
Error budget: the allowed failure rate (e.g., 0.1% of requests, or about 43 minutes of downtime per month)
If the error budget is exhausted, new features stop and reliability work takes priority
Post-mortems
Document what happened, what the impact was, and why it happened
Blameless culture: focus on systemic causes, not individual mistakes
Five Whys: ask “why” repeatedly to find the root cause rather than the proximate cause
Action items must be concrete, assigned, and time-bound
Circuit breaker implementation
Implement a CircuitBreaker class in Python or JavaScript that wraps a function.
It should track consecutive failures,
open after a configurable threshold,
fail fast while open,
and attempt recovery after a timeout.
Write tests that simulate a dependency that fails, recovers, and fails again,
verifying the breaker transitions through all three states correctly.
Single points of failure analysis
A diagram shows a web application with a load balancer,
two application servers,
one database,
and one external payment API.
Each component can fail independently.
Identify the single points of failure,
propose bulkhead and timeout strategies to limit the blast radius of each failure,
and discuss the cost (infrastructure, complexity) of each mitigation.
At what point does additional resilience add more complexity than it is worth?
Error path test coverage
A small data processing function is provided,
along with a test suite with 100% line coverage on the happy path.
Enumerate all error paths in the function
(e.g., invalid input, missing file, network failure, and malformed response).
Write tests for each error path using fault injection (i.e., mock the failing component).
Measure what fraction of error paths were untested before,
and explain why line coverage is a misleading metric for error-handling code.
Appendix: Exceptions
Exception hierarchies
Python: BaseException vs. Exception vs. specific types
KeyboardInterrupt and SystemExit inherit from BaseException, not Exception
JavaScript: Error base class plus built-in subtypes (TypeError, RangeError, SyntaxError, etc.)
Catching a parent class catches all subclasses
except Exception in Python does not catch KeyboardInterrupt
Custom exception types
Subclass the appropriate base: add fields for structured error data
Use custom exceptions to distinguish your errors from library errors
and to allow callers to catch specifically
Name exceptions as nouns describing the condition,
not the action (ConfigurationError, not FailedToLoadConfig)
Exception chaining
Python: raise NewException("context") from original_exception preserves the original traceback
Python: raise NewException("context") inside an except block implicitly chains
(visible as “During handling of the above exception, another exception occurred”)
Java: pass the original exception to the constructor of the new one (new RuntimeException("msg", cause))
Cleanup with finally and context managers
finally runs whether or not an exception was raised
Use for resource cleanup (closing files, releasing locks)
Context managers (with statement) encapsulate the try/finally pattern
Prefer them for resource management
__exit__ receives exception information and can suppress the exception by returning True
Do this rarely and intentionally
Anti-patterns
except Exception: pass: silently discards all errors
except Exception as e: print(e): visible but unactionable
Provides no context and does not propagate
Catching overly broad types: bare except: in Python catches KeyboardInterrupt and SystemExit
Raising Exception directly instead of a specific type: callers cannot catch it selectively
Checked vs. unchecked exceptions (Java)
Checked exceptions must be declared or caught
Unchecked (RuntimeException subclasses) need not be
Checked exceptions enforce handling at compile time
but lead to verbose, often-ignored throws declarations
Most modern languages (Python, C#, Kotlin, Swift) do not have checked exceptions
Would you like to have some real impact on the tech industry?
Do you have $100,000 to spend?
If you answered “yes” to both questions,
ask software engineering researchers
(the kinds of people who participated in It Will Never Work in Theory)
to design a study that companies could run internally
to measure the impact that genAI adoption by programmers is having on business outcomes.
Spend $50K to get expert reviews from both practitioners and (other) researchers,
publish all of the proposals with the reviews,
and award prizes of $25K, $15K, and $10K to the three best proposals.
(If you really want to have an impact,
do this in two rounds so that participants can hybridize their best ideas.)
$100K feels like a lot of money…
Really? Compared to what you’re spending on tokens?
Does anybody actually know how to measure genAI’s impact?
It’ll be interesting to find out.
(After all, “yes” and “no” are equally interesting answers.)
Will people write decent proposals for just a few thousand dollars?
No, but they’ll do it for the attention,
and for the chance to be involved in running the study if their proposal is a winner.
I’m currently making a few last changes to the third book in this series and trying to find an agent who will handle them. If you have middle-graders who would be interested in reading them and giving me feedback, please give me a shout.
Maddy Roo
Maddy Roo takes place in a world of anthropomorphic animals and patchwork robots. Its protagonist, Maddy, is a 12-year-old kangaroo whose younger sister, Sindy, is a “throwback” with no fur, scales, or tail. Their father was kidnapped by a raiding band of robots two years before the story opens; they and their mother have struggled to make ends meet since then.
While Maddy is out one evening with a goat boy named Gumption they rescue a damaged robot from a stream. Its regulator has been broken, which allows it to reveal that another raid is about to take place. Maddy and Gumption rush back to town to warn everyone. The raiders are driven off, but not before taking Maddy’s sister and two other children.
Maddy teams up with the rescued robot, Dockety, to get her sister back. The pair manage to catch up with the raiders and free the prisoners, but in the confusion that follows, Maddy, Sindy, and Dockety are stranded in a dangerous swamp called the Mire. They take refuge in an abandoned bunker, only to discover that it is the lair of a mad robot named Patient in Darkness, who is responsible for the raiding parties.
The trio escapes by bolting a flying suit onto Dockety and get back to town moments ahead of the raid. In the aftermath of the battle that follows, Maddy realizes that she knows how to free the bots that Patient has enslaved. She uses the flying suit to return to the bunker and break Patient’s control. As the story ends, Dockety reveals that Maddy’s father is still alive and is being held prisoner in the bot city of Heck.
In Heck
In Heck picks up several months after Maddy Roo. Dockety’s community of free bots has settled just outside Rusty Bridge, and Dockety has confirmed that Maddy’s father being held in the bot city of Heck. When Sindy accidentally activates a visiting Operator’s tech during a school demonstration and badly burns Special Leaf, the Operators insist on taking her to their headquarters in Sandy Bend, even though Special Leaf warns Maddy not to let them.
Maddy and Gumption stow away in the Operators’ wagon, but a rogue flying bot kidnaps Maddy mid-journey and delivers her to the mad bot Patient in Darkness. Patient claims the Operators in league with Central (the AI that controls Heck) which plans to exploit Sindy’s ability to activate Maker technology. Maddy escapes with a discombobulator that hides her from machines and a small cleaning bot she names Mouse, makes her way to Heck, and watches helplessly as the Operators hand Sindy over to Central’s bots.
Meanwhile, Gumption and Dockety seek help from a community of free bots in the forest. A reclusive bot called the Tailor disguises Gumption as a machine, and they reluctantly join forces with Patient. Inside Heck, Maddy finds her father, but is captured and placed in a virtual reality. There, Central reveals it is trapped by its own programming and longs to end.
The climax takes place in Central’s laboratory. Gumption’s disguise lets him briefly command Central’s bots, but Patient seizes control of Central through Sindy’s network connection. Sindy defeats Patient, Thoughtful turns on Special Blazes to save Sindy, and Dockety is nearly destroyed shielding Sindy from harm. The group escapes Heck with Maddy’s father and a handful of other prisoners. As the story ends Papa Roo dreamily announces that the Makers are awake and returning.
The Makers Return
The Makers Return picks up several months after In Heck. Special Leaf has died and left his house and his collection of ancient tech to Sindy. With Maddy and Gumption away in Sandy Bend, she is struggling to find her place in Rusty Bridge. She discovers a communicator that connects her to Violet, a young human girl aboard a failing spaceship called the Ark that has lost contact with Central and is running out of fuel. The ship’s commander, Captain Leung, decides it is time to return to the planet.
Special Blazes returns to Rusty Bridge with a new partner just as the Ark crash-lands in the swamp, where it is seized by a tentacled bot first encountered in Maddy Roo. Sindy uses her abilities to command the bot to release the ship, but a second attack creates chaos. Captain Leung seizes Sindy and, with Violet, flees in a shuttle. When the shuttle is forced down near an abandoned bunker, Patient in Darkness is waiting for them. The mad bot captures the group and uses a neural cap to read Captain Leung’s memories, confirming that the Makers have truly returned.
Violet discovers that Patient plans to use the Makers’ combat bots to destroy the ark. She escapes with Sindy and Mouse, and lead Patient (in a giant new body) back to Rusty Bridge. There, Violet wields a remote-control glove from Special Leaf’s hidden cache to disable Patient’s forces, and Mouse convinces the swamp creature to drag Patient under the water for good. In the conclusion, we see Violet and other children from the Ark settling into Rusty Bridge as Captain Leung calls the other arks home.