A chapter on mechanistic interpretability and the limits of AI explainability.
Abstract (español) ↗Philosophy PhD student at the University of Rochester working on the value alignment of AI systems and the interpretability tools to study them.
Lately: Just out of CAMBRIA with CBAI, where my capstone was on introspection fine-tuning: teaching a 1B-parameter model to detect and report perturbations in its own activation space. Finished BlueDot’s Frontier AI Governance course, and built an interactive map of who actually governs frontier AI. Now getting CNY AI Safety off the ground in Upstate New York. Our launch retreat is set for Nov 13–15 in the Finger Lakes, sixteen people, no prior AI safety background required.
— CNY AI Safety is supported by BlueDot Impact.
Writing
Papers, a book chapter, op-eds, and working notes.
Papers & Chapters
Working Papers
A philosophical look at how recommender systems bear on autonomy and human flourishing.
Manuscript ↗Media
On Large Reasoning Models and Apple’s “The Illusion of Thinking” — why a model’s displayed step-by-step “reasoning” isn’t evidence of actual reasoning, and language manipulation shouldn’t be taken to imply thought.
Clarín ↗El Litoral ↗Documenting my learnings across philosophy and AI.
Substack ↗Work
Projects, talks, teaching, and training.
Projects
A regional AI safety community started in Rochester, connecting students, researchers, and builders across Central and Western New York. Launched with a $7,000 BlueDot Impact Rapid Grant (June 2026); its first event is a launch retreat for sixteen people, Nov 13–15 in the Finger Lakes, open to newcomers as well as researchers. The longer-term goal is a corridor-style network of connected AI safety hubs across Rochester, Syracuse, Buffalo, and Ithaca/Cornell.
An interactive map of the institutions, actors, and mechanisms that govern frontier AI — organized by governance layer and stacked by how much real enforcement power each wields. Built as a thinking tool for BlueDot’s Frontier AI Governance course (June 2026).
Activation-based classifiers that measure and steer moral dispositions in language models.
Open-source framework for detecting algorithmic bias in LLMs using established audit-study methodologies.
Cosmos Institute grant project: a framework for continually learning virtue-theoretic signals in multi-agent systems.
Talks & Presentations
Teaching
Courses & Certifications
Honors & Awards
Curriculum Vitae
I work at the intersection of ethics, political theory, and artificial intelligence — studying virtue, wellbeing, and agency in AI systems, and building the interpretability tools to study them. As a Philosopher and Computer Scientist, I examine how algorithmic architectures and institutional norms co-determine the moral trajectories of AI. My long-term goal is to help build the conceptual and technical foundations for an ethically responsible AI ecosystem — one that aligns machine capability with human flourishing.
Research Areas
AOSPhilosophy of Artificial Intelligence · Ethics of Technology · Philosophy of Mind
AOIEpistemology · Causation · Identity & Persons · Phenomenology & Consciousness · Plato