Philosophy of AI · Interpretability

Philosophy PhD student at the University of Rochester working on the value alignment of AI systems and the interpretability tools to study them.

Lately: Leading three AI safety projects with nine mentees through SPAR. Hosting CNY AI Safety in Upstate New York, with support from BlueDot Impact. Two new posts on the blog.

Writing

Scholarly writing, and work for a wider audience.

From the blog

All posts

Papers & Chapters

2026
Abriendo la Caja Negra: Interpretabilidad Mecanicista y los Límites de la Explicabilidad de la Inteligencia Artificial
Opening the Black Box: Mechanistic Interpretability and the Limits of AI Explainability
Chapter in ‘Tratado de Inteligencia Artificial’ (ed. Fernando L. Depalma) · Editorial Hammurabi · forthcoming
Abstract (español) ↗
InterpretabilityPolicyEspañol
2025
Detecting and Steering LLMs’ Empathy in Action
arXiv:2511.16699

Studies empathy-in-action as a linear activation direction, and analyzes the detection–steering gap.

arXiv ↗Code ↗
InterpretabilityPhil. of AI

Manuscripts

2025
On Recommender Systems, Flourishing, and Autonomy
Unpublished manuscript

A philosophical look at how recommender systems bear on autonomy and human flourishing.

Manuscript ↗
Phil. of AIAutonomy

Media

2026
Pax Silica y la ruta de la seda digital
Pax Silica and the Digital Silk Road
Expert commentary · Diplomacia Activa (Argentina)

Consulted by Luca Nava for a piece on the US–China contest over AI infrastructure. I argue that open model weights buy a state sovereignty of inference but not of development: compute, semiconductors, cloud, and talent stay concentrated, so dependency migrates to less visible layers. Open weights are a necessary but not sufficient condition for autonomy — what they add is optionality.

Diplomacia Activa ↗
CommentaryEspañol
2025
La ilusión del razonamiento artificial
The Illusion of Artificial Reasoning
Op-ed · Clarín & El Litoral (Argentina)

On Large Reasoning Models and Apple’s “The Illusion of Thinking” — why a model’s displayed step-by-step “reasoning” isn’t evidence of actual reasoning, and language manipulation shouldn’t be taken to imply thought.

Clarín ↗El Litoral ↗
Op-edEspañol
2024–
Filosof.IA
Spanish-language blog & podcast

Documenting my learnings across philosophy and AI.

Substack ↗
Español

Projects

Current research, code, and talks.

Currently working on

Fall 2026
Constitutions and Reasons: virtue-based character training with reflect-update correction loops
Supervised Program for Alignment Research (SPAR) · Project lead

Open character training (Maiya et al. 2025) shapes model persona by fine-tuning on teacher demonstrations, but Anthropic’s production experience (“Teaching Claude Why,” 2026) found demonstrations alone insufficient: the gains came from teaching the reasons and identity behind behavior. We will build and test a virtue-based alternative on top of the OpenCharacterTraining infrastructure: excess/mean/deficiency contrastive data, rationale-annotated responses, and an iterated reflect-update correction loop that no current character-training work implements, evaluated head-to-head against the OCT baseline with ablations.

Project page ↗
AlignmentBehavioral evaluationPhil. of AI
Fall 2026
Disentangling persona vectors from emotion vectors in LLM activation space
Supervised Program for Alignment Research (SPAR) · Project lead

Persona vectors (Chen et al. 2025) and emotion representations (Sofroniew et al. 2026) have been studied independently, but plausibly overlap in activation space. We’ll measure their geometric and causal relationship to determine whether persona drift and emotional-state changes are mechanistically distinct failure modes; and whether interventions on one silently move the other.

Project page ↗
InterpretabilityBehavioral evaluation
Fall 2026
Does your assistant respect your agency? A behavioral benchmark for autonomy-preserving AI
Supervised Program for Alignment Research (SPAR) · Project lead

AI assistants constantly choose between empowering users and acting for them, and between honoring users’ stated goals and overriding them “for their own good.” We’ll build a systematic benchmark measuring whether models respect user agency, covering paternalism, manipulation, dependency-fostering, and value-substitution, with philosophically grounded rubrics and human-validated LLM-as-judge scoring.

Project page ↗
Behavioral evaluationSocietal impacts

Code & Sites

Talks & Presentations

Philosophical Problems of Algorithmic Agency
Escuela de Posgrado Newman — Perú
Slides (español) ↗
Creativity as Interplay
Rutgers–Columbia Undergraduate Philosophy Conference
Paper (PDF) ↗
The Why? Machine: Computing Causation
Iona University — Honors Thesis Day & Scholars Day
Video ↗

Curriculum Vitae

I work at the intersection of philosophy of mind and mechanistic interpretability, studying how virtue, wellbeing, and agency show up, and can be measured, in AI systems. My current research examines the epistemics of interpretability: what mechanistic explanations of model internals can and can't license us to claim.

Download PDF

Research Areas

AOSPhilosophy of Artificial Intelligence · Ethics of Technology · Philosophy of Mind

AOIEpistemology · Causation · Identity & Persons · Plato

Education

PhD in Philosophy
University of Rochester
BA in Philosophy & Computer Science
Iona University · Honors Program
Cybersecurity concentration · Magna Cum Laude.

Experience

Graduate Researcher
TRIAD, University of Rochester · with Jon Herington
Translational Research for Implementation of Augmented Decision-making. Three projects on verifying AI-written clinical notes: interpretability for medical AI, LLM-as-judge failure modes, and a deterministic dose-band checker built on 20M MIMIC-IV prescription records.
CAMBRIA Fellow
Selective 3-week ML research program for AI safety: mechanistic interpretability and RL, based on the ARENA curriculum.
Founder & Board Member
SALVé–Techlanna · New York
AI enablement for enterprises and government — including building systems for Tucumán’s legislature in Argentina.
Digital Consultant
Hopps · Freelance
Advised and built for clients on the Webflow platform — implementation, optimization, and training on best practices.
Webflow Developer
QRoom LLC · Contract, Remote
Built and maintained The Futur’s site and a custom dashboard for Pro members; produced brand assets in Figma and Adobe CC.
University Innovation Fellow
Stanford d.school (Hasso Plattner Institute of Design)
Selected for Stanford’s design-thinking fellowship; prototyped a campus-wide collaboration app.
Resident Assistant
Iona University · New Rochelle, NY
Residential life: enforced housing policy, counseled peers, ran educational programming, and mediated conflicts.
Fellow
Hynes Institute for Entrepreneurship & Innovation, Iona University
Ran digital campaigns, served as an institute ambassador, and led a campus 3D-printing initiative.
IT Specialist
Quadro Diseño y Comunicación · Buenos Aires, Argentina
Led an award-winning team on digital content campaigns reaching 6M+ impressions nationwide.

Publications

Media

Pax Silica y la ruta de la seda digital
Expert commentary · Diplomacia Activa (Argentina)
Consulted by Luca Nava for a piece on the US–China contest over AI infrastructure. I argue that open model weights buy a state sovereignty of inference but not of development: compute, semiconductors, cloud, and talent stay concentrated, so dependency migrates to less visible layers. Open weights are a necessary but not sufficient condition for autonomy — what they add is optionality.
La ilusión del razonamiento artificial
Op-ed · Clarín · El Litoral (Argentina)
On Large Reasoning Models and Apple’s “The Illusion of Thinking” — why a model’s displayed step-by-step “reasoning” isn’t evidence of actual reasoning, and language manipulation shouldn’t be taken to imply thought.
Filosof.IA
Blog & podcast · Substack
Documenting my learnings across philosophy and AI.

Selected Projects

A regional AI safety community started in Rochester, connecting students, researchers, and builders across Central and Western New York. Launched with a $7,000 BlueDot Impact Rapid Grant (June 2026); its first event is a launch retreat for sixteen people, Nov 13–15 in the Finger Lakes, open to newcomers as well as researchers. The longer-term goal is a corridor-style network of connected AI safety hubs across Rochester, Syracuse, Buffalo, and Ithaca/Cornell.
An interactive map of the institutions, actors, and mechanisms that govern frontier AI — organized by governance layer and stacked by how much real enforcement power each wields. Built as a thinking tool for BlueDot’s Frontier AI Governance course (June 2026).
Activation-based classifiers that measure and steer moral dispositions in language models.
Open-source framework for detecting algorithmic bias in LLMs using established audit-study methodologies.
Cosmos Institute grant project: a framework for continually learning virtue-theoretic signals in multi-agent systems.

Talks & Presentations

Philosophical Problems of Algorithmic Agency
Escuela de Posgrado Newman — Perú · Slides (español)
Creativity as Interplay
Rutgers–Columbia Undergraduate Philosophy Conference · Paper (PDF)
The Why? Machine: Computing Causation
Iona University — Honors Thesis Day & Scholars Day · Video

Teaching

Graduate Teaching Assistant
University of Rochester · Philosophy of AI; Ethics of Technology
Guest Lecturer
Philosophy of AI (PHIL 257), University of Rochester · “Argument Analysis”
Guest Lecturer
Ethics of Technology (PHIL 120), University of Rochester · with Randall Curren · “Autonomous Weapon Systems”
Guest Lecturer
Philosophy of Wellbeing (PHIL 239), University of Rochester · “Work and Play in the Shadow of AI”
AGI Strategy & Career Course Lead
Golem Lab · LATAM
Ad Honorem Teaching Assistant
Iona University · Logic; Philosophy of Psychology & Neuroscience

Certifications

Frontier AI Governance
BlueDot Impact · Credential
CS50AI: Introduction to Artificial Intelligence with Python
edX · HarvardX · Certificate · Credly

Honors & Awards

BlueDot Impact Rapid Grant — CNY AI Safety
BlueDot Impact · cnyaisafety.org
Stanford AI Learning Differences Hackathon — 1st Place
Stanford University
Cosmos Ventures Grantee
Cosmos Institute · Project
University Innovation Fellow
Stanford d.school · Directory
Dean’s Honors Scholarship
Iona University

Graduate Coursework

Selected Topics in AI Alignment
PHIL 595 · PhD Research · with Jens Kipper
Selected Topics in Philosophy of Mind
PHIL 591 · PhD Readings · with Grace Helton
Theory of Knowledge
PHIL 503 · with Earl Conee
Topics in Philosophy of Mind
PHIL 544 · with Grace Helton
Selected Topics in Philosophy of AI
PHIL 557 · with Rush Stewart
Selected Topics in Ancient Philosophy
PHIL 465 · with Randall Curren
Data, Algorithms, Justice
PHIL 435 · with Mark Povich
Selected Topics in Modern Philosophy
PHIL 470 · with Dante Dauksz
Selected Topics in Social and Political Philosophy
PHIL 523 · with Randall Curren
Seminar on Artificial General Intelligence
CSC 409 · audited · with Christopher Kanan
Metaphysics
PHIL 442 · with Paul Audi
Philosophy of Artificial Intelligence
PHIL 457 · with Jens Kipper