Ben Collier, PhD

Portfolio

Coding with AI Projects

Working projects I built with AI coding tools, several of them for 15-113 Effective Coding with AI. Each repo includes the prompts and build log, and I use them as examples in class.

Three iPhone screens from Ignatius at Home: the home page continuing the retreat's third day, Living Water, with Bloch's painting of the woman at the well; the guided prayer player reading the day's reflection, with its progress through the day's parts; and the retreat page with its days and the painting
Web app · Python, FastAPI, Claude Code, Claude, ElevenLabs, Supabase, Render · September 2026

Ignatius at Home

An app that turns retreat material, such as a handout, a few passages, or just an idea like "the parables of Jesus, seven days", into a guided audio retreat prayed one day at a time. Each day follows lectio divina: the passage read four times, a reflection, a closer look at the text's history and original language, silence between two bells, and a painting to pray with. I built it for 15-113, Effective Coding with AI, as the course's server-side project on Render.

  • The model plans each day and writes around the text, but never writes scripture. Scripture comes from the person's own material or the public-domain World English Bible.
  • It is free for anyone: open models on an academic cloud, free Microsoft voices, and the free tiers of six search services, with Claude and ElevenLabs as the paid tier.
  • Notes a person writes about themselves stay private, even in the build logs of the shared examples.
Rink Rivals comparing Sidney Crosby and Alex Ovechkin, with trophy badges and a season-by-season points chart
Web app · Python, Flask, Claude Code, the NHL API, OpenAI · September 2026

Rink Rivals

Type in two NHL players and see their careers side by side: every regular-season statistic with the leader highlighted, each player's draft position and trophies, and a season-by-season chart of their production. If you want it settled, a language model reads the whole comparison and picks a winner. It has to commit, explain the pick, and give the strongest argument against it. I built it for 15-113, Effective Coding with AI, as a study in what a small app owes the person using it.

Most of the interesting work was deciding what the app should refuse to do. The NHL's search endpoint does not rank results by relevance. A search for "Marc-Andre Fleury" returns Marc-Andre Dorion first, so taking the top result quietly shows the wrong career without any error. A scoring function fixes that, and a regression test checks that the endpoint still behaves this way. Skaters and goalies share almost no statistics, so the app declines to compare them rather than print numbers that are accurate and meaningless. Shots against and games played measure workload, not quality, so they are shown but never scored, which keeps a long career from passing for a great one. The OpenAI key stays on the server. GitHub Pages has no server to hold a secret, and GitHub Secrets are build-time variables that would end up in the published JavaScript.

This is the small, public version. I have also built a much larger hockey platform that is now being demonstrated to NCAA Division I programs for their coaches, managers, scouts, and players. Its code and login are private, so they are not linked here. It runs several coordinating agents and a data-analysis agent over a hockey database of more than 6 GB. I am happy to demo it in person for anyone interested in hockey data work at that scale.

Charlie the cocker spaniel, in a sweater, hopping across a voxel road in Charlie Road
Browser game · Kiro, Claude Code, three.js · September 2026

Charlie Road

A browser version of Crossy Road starring Charlie, our nine-year-old cocker spaniel, in place of the chicken. It keeps what makes the original work: the voxel look, one hop per input, endless roads, rivers, and railways, and a camera that punishes dawdling. Then it adds what matters to Charlie: tennis balls instead of coins, a bark that stops traffic, squirrels to chase, his real outfits as unlocks, a giant caterpillar toy worth five balls, and a flying saucer piloted by a squirrel where the eagle used to be.

The first version came from thirty minutes in class with Kiro. It went through five iterations, from colored rectangles to a fake-isometric board that never quite worked. I rebuilt it in an evening at home with Claude Code. I measured the camera angle from real gameplay screenshots and wrote a spec with acceptance criteria a machine could check, plus fairness rules for the world generator. Then I let an autonomous run build it one block at a time, checking each block by stepping the simulation through a debug harness instead of trusting a screenshot. After I played it, a second round added the bark, the saucer, combos, a daily challenge, and a shareable trading card. Every prompt, decision, and build step is in the repo.

A 3x3 Raven's Progressive Matrices puzzle with the last cell missing and six answer options
Experiment · Python, Claude Code, 16 language models · September 2026

Can a computer pass an IQ test?

  • 13random guessing
  • 34a student's 2017 program
  • 59expert system, no language model
  • 93current language model

Three programs, one from 2017 and two from this year, take the same 96 Raven's Progressive Matrices, the wordless reasoning test from 1936 that still shows up in hiring. A current language model, given the puzzle sheet and every cell as images, nearly aces it.

Sixteen language models from nine companies ran the same puzzles with the same prompt, so the write-up compares accuracy, cost, and speed directly. Four models tied at 92 while their prices differed by almost five to one. Five scored below the program with no AI in it. One release changed how the model works at answer time, and that moved its score more than two years of scaling did on either side. A neural network trained on thousands of invented puzzles scored 10, worse than guessing, and the explanation is the most useful lesson in the repo. A script generates every number from the raw results, and the prompt history quotes every instruction that produced them.