How I'm Watching WWDC26 With Codex
WWDC week has always been a funny mix of excitement and overload for me.
There are the big moments from the keynote, the more technical State of the Union, the individual sessions, the screenshots people share, the tweets with tiny details, and then my own messy reactions as I notice what actually matters to me.
It’s a lot, and it is not the first time I write down how I try to survive that flood. Back in 2015 I wrote about watching WWDC videos with a list, Evernote, screenshots, and a lot of 2x playback. A few years later I went deep on what I wanted from a notes app, and that search eventually led me to Obsidian, which genuinely changed my life. In 2020, when Apple moved to the new remote format, I changed the workflow again and started being more selective with what was important enough to write down. And after that I kept publishing the raw, personal versions of the notes, like WWDC21 and WWDC23.
The pattern is always the same: I want the notes to preserve what caught my attention, not pretend to be a complete archive of the conference. Even though now we have AI, that main goal still hasn’t changed.
This year I’m trying a slightly different setup: YouTube on one side, Codex on the other, both full screen.

It sounds simple, but it is working surprisingly well, mainly because Codex is a beast of an app that removes a lot of the friction around keeping the useful bits around.
Dumb note 😂 I have Codex on the left even though I prefer it on the right, because otherwise every time I mouse over its left edge, the conversation sidebar shows up, blocking me from actually seeing my conversation. I love the app, and the speed software moves nowadays with AI, but the amount of bugs that all software has due to that is a bit annoying.
Starting with the transcript
The first thing I do is give Codex the YouTube URL for the session I’m watching. Codex then uses a little YouTube notes skill I’ve been using with my Obsidian vault for a while. It downloads the transcript and keeps it around for the rest of the conversation.
I still truly believe having your own PKM system, writing yourself, and learning are very important. That’s why the transcript matters, but not because I want an AI summary of the whole video or a giant dump of everything Apple said. The transcript is there as a supporting source for my own input.
So I ask Codex to either create or find the relevant note in my Obsidian vault. In this case, it found the WWDC26 note and we agreed on a very specific workflow:
I watch the video myself. When I mention something, Codex can go back to the transcript, find the surrounding explanation, and complement my note with the missing bits. The important constraint is that it should support my words, not replace them or add walls of text.
Watching normally
With the context set up in a Codex conversation, I just watch.
When something jumps out, I leave a message. Sometimes it is short and conversational, almost like a tweet: “Dynamic Profiles sounds very cool”. Other times I pause the video and dump my thoughts, future things I need to check, or tasks to do.
Codex turns those messages into useful notes by checking the transcript. For example, when I noticed server-side models support, it added that the Foundation Models framework can call models such as Claude, Gemini, and others for more complex workflows, including tool calling and guided generation. When I wondered which developers get no-cost Private Cloud Compute access, because I didn’t catch it in the video, it found the exact line: developers with fewer than 2 million first-time App Store downloads.
Something that Codex does very well is enqueuing messages. It’s probably the AI app that does this the best, or at least it was when I started using it. It means it doesn’t matter if I write longer thoughts or a burst of new messages. Codex will do the work and take them when ready. At the start I thought about using faster modes, but because enqueuing works so well, there is no rush. I can keep watching the video while the agent does its work.
That is the part I like. I get to stay in the flow of watching, while the assistant does the tiny retrieval work that would otherwise make me pause, rewind, search, and lose the thread.
Screenshots and Appshots as attention markers
The other nice part is screenshots, or even better, Appshots!
When something visual appears in the video, I can throw the screenshot into the Codex conversation. That screenshot is not just an image to save. It is a signal: this is the thing I care about right now.
Codex Appshots makes this even faster. I only have to press both Command keys and the screenshot is automatically inserted into the Codex input field, with extra context from Chrome. I can do that several times during a segment, then when the presenter finishes that part I add my own thoughts and send everything together.
For example, Apple showed a layered visual of Apple Intelligence that I really liked. That visual was a useful mental model, so I sent it over. Codex attached it to my WWDC26 note following my vault rules and added bullet points describing the architecture.

The screenshot is a starting point. Sometimes it gets included in the final note, but other times it is just a visual reference to pull the final details from.
Side channels
It also works for the side WWDC conversations happening on socials, which is often the hardest part to track and where more interesting tidbits surface.
Now I just need to paste the tweet URLs into the thread. Codex can inspect them and decide where they belong. Not every tweet is about the exact video I’m watching, and that is fine. It will add them to the appropriate place in the note or my vault, and it will even note down things for me to check later.
That is another small but important difference from dumping links into a note. The assistant can put the link where it actually belongs.
Why this works
The setup works because the assistant has the boring context loaded and I keep control over what matters. With Codex I don’t have to think too much about keeping the conversation tidy. The app does tremendous magic with auto-compaction, so you can have single long threads without worrying about managing them. People highlight this as something great for coding, but I’ve found it more useful for tasks like this, where I want to focus on learning and my thoughts.
And my thoughts are what matter. I don’t want an exhaustive WWDC archive. Apple already has that. I want my WWDC notes: the details that make me stop the video, the APIs I want to inspect later, the screenshots that explain a concept better than a paragraph, and the small questions I ask while watching.
The transcript makes the assistant useful and the Obsidian note gives us a persistent place to put everything, which we can tweak and iterate later. The conversation gives me a low-friction input surface. And the instructions make it behave as I want it to: not summarize everything, but instead help me remember what I already noticed.
I’m building my PKM, improving myself, not making a global Wikipedia. That’s often the huge mistake that note takers make, especially in the age of AI.
That is the version of AI assistance I keep finding valuable. Not a replacement for attention, but a way to preserve it.