Lectures
Part 1: Foundations of Coding Agents
Lecture: Course Introduction + How LLMs Actually Work
This lecture introduces agentic software engineering and explains the LLM machinery beneath coding agents: tokenization, next-token prediction, transformer basics, model training, context windows, and statelessness. Understanding these foundations clarifies how agents can produce code, but also why agent output can be nondeterministic or confidently wrong—and why deliberate context management and verification are essential throughout the course.
- Lecture Notes (md)
- Video of in-class demo (mpg)
- Readings/Videos to cover before next lecture
- [required] Karpathy, Intro to Large Language Models
- [recommended] 3Blue1Brown, Transformers, the tech behind LLMs and Attention in transformers, step-by-step — additional shorter material with good visuals
- [optional deeper dive into how LLMs work] Karpathy, Deep Dive into LLMs like ChatGPT; The original paper on attention/transformers by Vaswani et al., Attention Is All You Need (skim §1–2); A key paper on Training language models to follow instructions with human feedback (RLHF) - InstructGPT - Training language models to follow instructions with human feedback
Lecture: From LLM to Agent: The Loop, System Prompts, and Tool Calls
This lecture explains the concept of a agent “harness” - software that wraps an LLM (model) and turns it into an agent by maintaining message history, defining tools, executing tool requests, and repeating the model–tool loop. This model-versus-harness distinction helps us reason clearly about agent capabilities, safety boundaries, and failure modes throughout the course.
- Lecture Notes (md)
- Readings/Videos to cover before next lecture
- [required] The Carbon Layer - YouTube Channel - Harness Engineering Masterclass: Technical Deep Dive on how to build Agentic Systems — up to timestamp 14:00 is good enough for preparing for L03. This is a Week 1-3 core reference (our toy agent will address the first several “primitives” (building blocks) for a coding agent that are introduced in the video.
- [optional – see how Claude handles tool calls] Anthropic API docs, Tool use overview (skim to recognize the basic ideas)
- [recommended] Yao et al., ReAct (skim §1–3).
- Exercise: Ex. 1 — transcript critique (md), due before Lecture 04.
Lecture: Claude Code Hands-On: Permissions, CLAUDE.md, and Plan Mode
This lecture describe how what one sees in Claude Code’s “industrial strength” interface relates to the concepts in the “toy agent” harness and loop actions presented in the previous lecture. It describes the basic information and commands that a developer provides to the agent to control the production of code: permissions, project instructions, explicit context, reusable commands, and plan mode. This introduction establishes the habits developers need to work with coding agents deliberately and safely. It also provides a first understanding of how developers control the information that feeds into the model context, i.e., exactly what content is being supplied as input to the LLM to get it to produce code.
- Lecture Notes (md)
- Annotated transcript of the in-class demos (md) — a real session,
/initthrough plan mode, with commentary. Includes instructions for reproducing it yourself.- Companion: the same prompt run three ways, with and without CLAUDE.md and plan mode
- Starter project used in the demos:
tictactoe-starter.tar.gz— extract it and run the tests first; you should see 90 passing
- Readings/Videos to cover before next lecture
- [required] The Carbon Layer, Harness Engineering Masterclass: Technical Deep Dive on how to build Agentic Systems — finish the video. You watched to 14:00 for this lecture; the next lecture looks at harnesses other than Claude Code, and the rest of this video is the vocabulary for it.
- [recommended] Claude Code docs on permission modes, memory/CLAUDE.md, and plan mode — curated links are in the course materials’
technical-concepts.md.
- Exercise: Ex. 2 — codebase comprehension (md), due before Lecture 05.
Lecture: Other LLMs and Harnesses
A discussion session on coding agents other than Claude Code — Codex, OpenCode, and Grok among them — with live demos. The comparison runs along the axes that matter once you are actually working: whether the harness is source-open, whether it shows you what is in your context and what you have spent, what it does when the conversation is compacted, and how portable your accumulated project context is if you move between tools.
A suggestion, not an assignment: bring a laptop if you’d like to try a few models yourself through OpenCode. Its free tier costs nothing, so you can form your own opinion without spending plan allowance, money, or a subscription commitment. Worth doing before Exercise 4 either way.
- No lecture notes for this session — it is discussion and live demonstration.
Lecture: Prompting + Spec-Driven Development
Rescheduled to a later week.
Lecture 4 uses contrasting development attempts to show how precise requirements, durable specifications, existing patterns, and built-in verification matter more than clever prompts. The concept of Spec (Specification)-driven development is introduced: organizing the development process are declarative specifications of what should be built, the desired behavior of the system, and how the system should be verified (rather than focusing on the details of coding and how system behavior is implemented). Spec-driven development is central to the agentic development because it expresses the developer’s intentions in human-reviewable documents while also directing agents with a stable source of truth as projects grow.
- Lecture Notes (md)
Lecture: Anatomy of a Coding Agent: The Basic Building Blocks
Having understood what an LLM is, the basic idea of a chatbot, the basic idea of a coding agent, and having done initial experimentation with Claude Code, this lecture repositions the introductory material onto a more rigorous footing by introducing architectural building blocks for coding agent implementations (our source material uses the term “primitives” instead of building blocks – we will use these terms interchangeably in this lecture). These building blocks represent concepts that a coding agent harness needs to implement. Identifying distinct harness resposibilities as building blocks is a good pedagoical tool – it is enables us to understand (and even build) a harness in bite-sized chunks. It also clarifies for us the distinct roles that are played by the code that you may see in a harness.
We’ll introduce these building blocks and tie the lectures from the last two weeks together by discussing how the concepts are realized in both the toy coding agent and in Claude Code.
This lecture also sets the groundwork for the exercise on building a simple coding agent in Python. Building this small agent demystifies production coding tools like Claude Code and makes their key reliability and safety mechanisms concrete enough to inspect, test, and improve.
Lecture: Context, Cost, Verification, and the Road Ahead
Lecture 6 examines the cost and quality consequences of growing context, practical strategies for managing sessions, and the verification gates needed to catch plausible but incorrect output. It brings the foundations unit together around the course’s larger goal: producing trustworthy software with agents as project scale and autonomy increase.
- Lecture Notes (md)
Part 2: Agentic Software Engineering Principles
Lecture: Specifications, Realizations, and Conformance
This lecture introduces the four terms the rest of the module is built on — specification, realization, conformance, and verification — using a running example of a folder of markdown study notes governed by a written format specification. It presents the five quality properties a specification should have (unambiguous, internally consistent, externally consistent, complete for its level of abstraction, traceable) and then turns those properties from developer guidelines into an audit that an agent can actually run. The lecture distinguishes three kinds of verifier — human, agent, and algorithmic — compares them on cost and repeatability, and closes with six ways a verification result can be invalid even when the check reports success. A theme that recurs throughout the module appears here: in agentic development, more of the developer’s time goes into writing specifications and setting up the development process, and less into writing the code.
- Lecture Notes (md)
- Handout: the note-set starter files — the draft specification, the loader
CLAUDE.md, and the three process documents (md) (pdf) - Demo: the note-set specification demo, part 1 (md)
- Demo material, including the starter tree the demo begins from (folder)
- Exercise: SE Module Ex. 1 — conformance three ways (md).
Lecture: The Concept of Operations, Operations with Contracts, and the Specification as an Invariant
This lecture introduces the concept of operations (ConOps) — a specification of purpose that says what a system is for, who uses it, and which operations they perform, while saying nothing about how it is built. Each operation is written as a contract with a precondition and a postcondition, applying Meyer’s design-by-contract to the work an agent performs rather than to a method call. The lecture then presents conformance as an invariant rather than a milestone: it is re-established after every operation and after every amendment. A live change of requirements is carried through the specification, the verifier, the derived operations document, and every affected note in a single commit. It closes with five claims about specifications that answer the common assumption that one can write a specification, hand it to an agent, and receive a finished implementation.
- Lecture Notes (md)
- Handout: the concept of operations and the process documents this lecture adds (md) (pdf)
- Demo: the note-set specification demo, part 2 (md)
- Companion: a prompt template for planning and realizing a system from specifications, with a filled example (md)
- Exercise: SE Module Ex. 2 — operations under contract (md).
Lecture: Specifying a Game: One System, Several Specifications
This lecture moves the method from markdown notes to code — a command-line tic-tac-toe on a nine-by-nine grid where five in a row wins. Three things change: the specification begins as an informal sketch rather than a draft, the realization is a program, and one program turns out to need not a single specification but a family of them — concept of operations, rules of play, display format, input grammar, interaction flow, interface contract, example session, and verification obligations — each fixing what the others leave open. The lecture audits the sketch, produces a gap list of fourteen open items, and shows the developer ruling on each and recording the ruling as an amendment with a version bump. The larger point is that in agentic development the developer’s attention shifts from writing code to designing the development framework: choosing which kinds of specification to use, writing the audit rules, and defining what conformance means. If nobody decides those things, the agent decides them by default, and the decision surfaces later in the shape of the code.
- Lecture Notes (md)
- Demo: the game demo, part 1 — auditing the sketch and deriving the specifications (md)
- Demo material, including the starter tree (folder)
- Exercise: SE Module Ex. 3 — the game-demo audit (optional) — to be posted.
Lecture: Verifying a Game: Tests as Executable Specification
This lecture treats a test suite as a specification that a program can check: every test cites the clause it verifies, may assert more tightly than that clause but never more loosely, and the suite together with a branch-coverage threshold becomes a single gate the developer runs. The centerpiece breaks the pairing of tests and code three ways — deleting a test, introducing a subtle indexing bug, and writing wrong code together with a matching wrong expectation — and asks after each what the green or red gate did and did not establish. The conclusion is that coverage is a property of the tests and the code, not of the tests and the specification: it reports which code the suite’s claims reach, not whether those claims are right. The lecture closes with fixture files as language-independent executable specifications that are never edited to make a test pass, and that are obliged to state their own blind spots.
- Lecture Notes (md)
- Demo: the game demo, part 2 — the test suite and the coverage gate (md)
- Reference implementation, with the fixtures and the verified suite (folder)
Lecture: Reversi Kickoff: Applying the Method
This lecture recaps the module’s method as one chain — from an idea to a sketch, a concept of operations, a specification, plans, code, and a verified program — and launches Project 1, which takes a text-based Reversi through that chain. Students play a Reversi demo, read the client’s concept-of-operations sketch, and sort what it states, implies, and leaves out. The sketch carries a seeded tension: it says the game ends when the board is full, and that a player with no legal move forfeits the turn — together they leave undecided two stuck players with squares still empty, which a nine-move wipeout shows can happen. The rest of the lecture walks the project brief step by step, showing each step’s prompt and what finished looks like on the tic-tac-toe reference; the Reversi gap list, rulings, specification, and code are the student’s own work.
- Lecture Notes (md)
- Demo: the Reversi walkthrough (md)
- Project: Project 1 — Reversi, specification-first, assigned at this lecture, due Thursday, October 8 (md)
- Starter: the client’s concept-of-operations sketch, the loader
CLAUDE.md, the process documents, and the test gate’s configuration — noCONOPS.md,SPECS.md, code, or tests yet; writing those is the project (folder) (download)
- Starter: the client’s concept-of-operations sketch, the loader
Lecture: Agent Skills
This lecture turns a method into something an agent can follow on request: a skill, a folder with a SKILL.md whose name and description the agent sees at the start of every session and whose body it loads only when the skill is invoked. It covers what a skill is and when to write one instead of a CLAUDE.md/AGENTS.md rule, a hook, or a subagent; how the harness finds, lists, loads, and invokes skills, and how to stop one from firing on its own; where skills live for Claude Code (.claude/skills) and for other agents (.agents/skills); and what makes a skill good. Two published skills show how skills evolve and how their authors evaluate them. The worked example is a skill that turns a concept-of-operations sketch into a gap list and waits for rulings, built before class and evaluated on the tic-tac-toe sketch and on a sketch it was never tuned on; its scores show why a skill must be tested on something it was not written against.
- Lecture Notes (md) · Slides (md)
- Worked-example skills: v1
conopsand v2conops-from-sketch(folder) - Exercise: build a skill that finds what a
SPECS.mdmust settle in a concept of operations, and evaluate it; due Tuesday, October 13 (md)- Starter: tic-tac-toe’s process documents and reference
CONOPS.md(folder) (download)
- Starter: tic-tac-toe’s process documents and reference
Lecture: MCP — The Tool Contract as a Specification
This lecture adds the last kind of specification the module introduces: the tool contract. The concept of operations names a computer opponent and says nothing about how it is built, and an LLM can fill that role through a Model Context Protocol server, where a tool’s type hints become its schema and its docstring is the entire interface the model can read. The lecture builds and registers a server for the game, then deliberately sabotages the docstring to show that the contract is load-bearing: a description that misstates how the board is indexed changes how the agent plays. It also treats a tool-call transcript as evidence against a concept-of-operations scenario.
- Lecture notes to be posted.
Part 3: Agentic Development at Scale
Lecture material to be announced.