// HACKER NEWS — CYBERSECURITY
What I did at Recurse Center
I spent the summer at Recurse Center, a programming retreat in Brooklyn, at the recommendation of my friend Cory. These are some things I did there.
RC has regular study groups that anyone can organize and schedule. For example, people would get together every Friday to work through old Advent of Code challenges together. I did Agentic Adventures, Practical Deep Learning, Math Monday, and a short series investigating an open source board game AI.
Agentic Adventures is a discussion group for people interested in modern LLMs and agents; we did different things every week. Some highlights:
This was a study group working through the (first half of the) book Practical Deep Learning by Jeremy Howard and Sylvain Gugger; we worked our way up through classical machine learning approaches, used or adapted prebuilt models, and then learned how to build up actual neural network models. The book gave a high level practitioner’s introduction, which was good for building intuition and having a sense for how they work, as well as what can be built with them. Since I’m not interested in research myself, this was the level of understanding I wanted.
In a coffee chat, Sophia and I discovered we shared a passion for math and decided to start a discussion group for mathematical topics. This was meant to be a fun distraction and chance to learn and explore, so we would spend maybe the first half of each meeting on learning and then split into pairs to build things based on what we learned. Some fun things we worked on were:
Tony started a short sequence to dig into the code of the Keldon AI for playing the board game Race for the Galaxy. We spent the first session just learning the rules and playing the AI, then dug into the code - an old school two-layer neural network with curated features. Claude was pretty good at interpreting the weights of the neural network and explaining what it reacted to; we spent some time trying to understand what the Keldon AI had learned about strategy. Takeaway was that econ strategies are much stronger than military in the base game (coinciding with player sentiment). One of the things that stood out to me was how many of the strongest weighted nodes corresponded to the presence of individual cards.
I wrapped up the mini language, dodo, that I had made as part of my application to RC. This included what was probably the most interesting part, writing the tricky recursive code to implement the match statement. This was an interesting balance with LLMs - I didn’t let them write any of the code for me, but it was nice to have someone writing up a nice spec and writing test cases and example programs. LLMs also made it easy to share what I did - try it out in a vibe-coded online REPL here!
RC hosts a lot of talks by people discussing their projects. One I attended early in batch was by Josh on the ZIP file format, and his investigations about inconsistencies between implementations. It dove deep into details that I’d never thought about before. For example, if the unzipper that your virus scanner uses handles edge cases slightly differently than the unzipper that you use, it might see a file with malware in it as innocent, because it doesn’t unzip the virus. This is a security vulnerability!
One part of the talk that caught my attention was the digression on the compression algorithm that zip uses, called DEFLATE. Kevan and I thought that reimplementing it sounded like a fun exercise, so we paired on it in a few sessions over the next few weeks.
DEFLATE is basically LZ77 (generalized run length encoding) + Huffman codes. It compresses text (or actually any sequence of bytes, but here we consider it as text) in two ways: (1) replacing repeated sequences with backreferences, e.g. “to be or not to be” could become “to be or not (13, 5)” where the part in parentheses means “go back 13 bytes, then copy 5 bytes from there”, and (2) Huffman encoding, where we assign shorter sequences of bits to represent common bytes,