Joe Kwon

Joe Kwon

Trying to help steer towards better (AI-entangled) futures!

AI is poised to be deeply transformative. I think about how to make that go well!

work

trajectory

Astra Fellowship

I worked as a strategy fellow with Tom Davidson (Forethought) and Fabien Roger (Anthropic) on secretly loyal AI: the risk that an AI system could be deliberately trained to appear aligned with an institution's goals while covertly serving a different actor's interests. I focused on threat modeling and designing ML experiments that stress-test this scenario and support the broader research agenda.

Center for AI Policy and Centre for the Governance of AI

In 2025 I moved to DC to work on AI policy and governance: first at the Center for AI Policy, writing reports on AI agents, cybersecurity, and autonomous systems, then GovAI's DC fellowship, working on risks from internal AI deployment and metrics for tracking automated AI R&D. This was refreshing because the questions felt immediately important and impactful. I enjoyed communicating ideas and recommendations to people (tens of thousands read my reports in total), and it led to being invited as a panelist on a Georgetown × World Bank conference on "Making AI Work: What Firms and Workers Need."

activation steering

I worked with David Krueger's group testing activation steering methods. At the time it was unclear how well these techniques actually worked or where they broke down. We compared bottom-up approaches like function vectors against top-down ones like in-context vectors on shared in-context learning tasks, and found they have different strengths: in-context vectors are better at broad behavioral shifts, function vectors at tasks demanding precision.

LG AI Research

In 2023 I was a research engineer on multi-lingual LLMs under Honglak Lee, working with Lajanugen Logeswaran, Dongsub Shim, and Tolga Ergen. Synthetic data, pretraining, finetuning, evals. One thread I liked: leveraging language-invariant concepts so models can learn new languages more efficiently.

MIT

After college I joined Josh Tenenbaum's Computational Cognitive Science Lab, working closely with Sydney Levine on moral cognition: how people reason about rules, norms, and each other. We built models that tried to capture the structure of moral judgment, something I think matters for AI alignment too. Separately, I worked with Stephen Casper and Dylan Hadfield-Menell on red-teaming methods for systematically finding where language models fail.

early AI safety

Around 2020 I started paying attention to the surprising capabilities emerging in AI systems. I worked on one of OpenAI's early RLHF projects under Long Ouyang and Jeff Wu. It was my first hands-on experience with LLMs, and it got me scaling pilled. Then at Berkeley with Jacob Steinhardt and Dan Hendrycks, I worked on out-of-distribution detection, AI forecasting, and building evaluations for ML systems.

Yale

Studied CS and psychology. The summer before sophomore year, I worked with Gabriel Kreiman and Mengmi Zhang at Harvard/MIT Center for Brains, Minds, and Machines on visual cognition and context reasoning. It was my first research experience and I'm grateful they invested their time in a mostly floundering freshman. During school I worked in Julian Jara-Ettinger's lab, building computational models of social cognition.

pre-college

Mostly spent my time hanging out with friends and consuming a ton of content online. I was pretty directionless: no real sense of what I ultimately cared about or wanted to pursue, just chasing whatever felt good in the moment, not anchored to any ideals. But experiencing CTY and Canada/USA Mathcamp was special and invigorating. They were the first environments where I felt intellectually excited about ideas and the people around me.

rabbit holes

reading
  • The Gentle Romance: Stories of AI and humanity — Richard Ngo
  • The Night Circus — Erin Morgenstern
  • The Book of Five Rings — Miyamoto Musashi
listening
hip hop
I LAY DOWN MY LIFE FOR YOU JPEGMAFIA experimental / industrial
LP! (Offline) JPEGMAFIA experimental / glitch
jazz(y)
The Black Saint and the Sinner Lady Charles Mingus avant-garde
Hot Rats Frank Zappa jazz-rock
art pop
LUX Rosalía orchestral
La Vida Era Más Corta Milo j contemporary folk
Vanisher, Horizon Scraper Quadeca folktronica
electronic
I Love My Computer Ninajirachi house / dance / pop
Allbarone Baxter Dury synth pop / electropop
The Provocateur ADÉLA pop / dance / house
rock
Fetch Melt-Banana noise / experimental
Pain to Power Maruja post-punk / jazz
looking

Updating soon.

bookmarks