Philosophy PhD student at the University of Rochester working on the value alignment of AI systems and the interpretability tools to study them.
Lately: Leading three AI safety projects with nine mentees through SPAR. Hosting CNY AI Safety in Upstate New York, with support from BlueDot Impact. Two new posts on the blog.
Writing
Scholarly writing, and work for a wider audience.
From the blog
All posts
The paradox of the just city and its first philosopher-king.

Why do Plato’s degenerate regimes fall apart?
Papers & Chapters
Manuscripts
A philosophical look at how recommender systems bear on autonomy and human flourishing.
Manuscript ↗Media
Consulted by Luca Nava for a piece on the US–China contest over AI infrastructure. I argue that open model weights buy a state sovereignty of inference but not of development: compute, semiconductors, cloud, and talent stay concentrated, so dependency migrates to less visible layers. Open weights are a necessary but not sufficient condition for autonomy — what they add is optionality.
Diplomacia Activa ↗On Large Reasoning Models and Apple’s “The Illusion of Thinking” — why a model’s displayed step-by-step “reasoning” isn’t evidence of actual reasoning, and language manipulation shouldn’t be taken to imply thought.
Clarín ↗El Litoral ↗Documenting my learnings across philosophy and AI.
Substack ↗Projects
Current research, code, and talks.
Currently working on
Open character training (Maiya et al. 2025) shapes model persona by fine-tuning on teacher demonstrations, but Anthropic’s production experience (“Teaching Claude Why,” 2026) found demonstrations alone insufficient: the gains came from teaching the reasons and identity behind behavior. We will build and test a virtue-based alternative on top of the OpenCharacterTraining infrastructure: excess/mean/deficiency contrastive data, rationale-annotated responses, and an iterated reflect-update correction loop that no current character-training work implements, evaluated head-to-head against the OCT baseline with ablations.
Project page ↗Persona vectors (Chen et al. 2025) and emotion representations (Sofroniew et al. 2026) have been studied independently, but plausibly overlap in activation space. We’ll measure their geometric and causal relationship to determine whether persona drift and emotional-state changes are mechanistically distinct failure modes; and whether interventions on one silently move the other.
Project page ↗AI assistants constantly choose between empowering users and acting for them, and between honoring users’ stated goals and overriding them “for their own good.” We’ll build a systematic benchmark measuring whether models respect user agency, covering paternalism, manipulation, dependency-fostering, and value-substitution, with philosophically grounded rubrics and human-validated LLM-as-judge scoring.
Project page ↗Code & Sites
A regional AI safety community started in Rochester, connecting students, researchers, and builders across Central and Western New York. Launched with a $7,000 BlueDot Impact Rapid Grant (June 2026); its first event is a launch retreat for sixteen people, Nov 13–15 in the Finger Lakes, open to newcomers as well as researchers. The longer-term goal is a corridor-style network of connected AI safety hubs across Rochester, Syracuse, Buffalo, and Ithaca/Cornell.
An interactive map of the institutions, actors, and mechanisms that govern frontier AI — organized by governance layer and stacked by how much real enforcement power each wields. Built as a thinking tool for BlueDot’s Frontier AI Governance course (June 2026).
Activation-based classifiers that measure and steer moral dispositions in language models.
Open-source framework for detecting algorithmic bias in LLMs using established audit-study methodologies.
Cosmos Institute grant project: a framework for continually learning virtue-theoretic signals in multi-agent systems.
Talks & Presentations
Curriculum Vitae
I work at the intersection of philosophy of mind and mechanistic interpretability, studying how virtue, wellbeing, and agency show up, and can be measured, in AI systems. My current research examines the epistemics of interpretability: what mechanistic explanations of model internals can and can't license us to claim.
Research Areas
AOSPhilosophy of Artificial Intelligence · Ethics of Technology · Philosophy of Mind
AOIEpistemology · Causation · Identity & Persons · Plato