Philosophy of AI · Interpretability

Philosophy PhD student at the University of Rochester working on the value alignment of AI systems and the interpretability tools to study them.

Lately: Just out of CAMBRIA with CBAI, where my capstone was on introspection fine-tuning: teaching a 1B-parameter model to detect and report perturbations in its own activation space. Finished BlueDot’s Frontier AI Governance course, and built an interactive map of who actually governs frontier AI. Now getting CNY AI Safety off the ground in Upstate New York. Our launch retreat is set for Nov 13–15 in the Finger Lakes, sixteen people, no prior AI safety background required.

— CNY AI Safety is supported by BlueDot Impact.

Writing

Papers, a book chapter, op-eds, and working notes.

Papers & Chapters

2026
Abriendo la Caja Negra: Interpretabilidad Mecanicista y los Límites de la Explicabilidad de la Inteligencia Artificial
Opening the Black Box: Mechanistic Interpretability and the Limits of AI Explainability
Chapter in ‘Tratado de Inteligencia Artificial’ (ed. Fernando L. Depalma) · Editorial Hammurabi · forthcoming

A chapter on mechanistic interpretability and the limits of AI explainability.

Abstract (español) ↗
InterpretabilityPolicyEspañol
2025
Detecting and Steering LLMs’ Empathy in Action
arXiv:2511.16699

Studies empathy-in-action as a linear activation direction, and analyzes the detection–steering gap.

arXiv ↗Code ↗
InterpretabilityPhil. of AI

Working Papers

2025
On Recommender Systems, Flourishing, and Autonomy
Unpublished manuscript

A philosophical look at how recommender systems bear on autonomy and human flourishing.

Manuscript ↗
Phil. of AIAutonomy

Media

2025
La ilusión del razonamiento artificial
The Illusion of Artificial Reasoning
Op-ed · Clarín & El Litoral (Argentina)

On Large Reasoning Models and Apple’s “The Illusion of Thinking” — why a model’s displayed step-by-step “reasoning” isn’t evidence of actual reasoning, and language manipulation shouldn’t be taken to imply thought.

Clarín ↗El Litoral ↗
Op-edEspañol
2024–
Filosof.IA
Spanish-language blog & podcast

Documenting my learnings across philosophy and AI.

Substack ↗
Español

Work

Projects, talks, teaching, and training.

Projects

Talks & Presentations

Philosophical Problems of Algorithmic Agency
Escuela de Posgrado Newman — Perú
Slides (español) ↗
Creativity as Interplay
Rutgers–Columbia Undergraduate Philosophy Conference
Paper (PDF) ↗
The Why? Machine: Computing Causation
Iona University — Honors Thesis Day & Scholars Day
Video ↗

Teaching

Graduate Teaching Assistant
University of Rochester · Philosophy of AI; Ethics of Technology
Guest Lecturer
Ethics of Technology (PHIL 120), University of Rochester · with Randall Curren · “Autonomous Weapon Systems”
Guest Lecturer
Philosophy of Wellbeing (PHIL 239), University of Rochester · “Work and Play in the Shadow of AI”
AGI Strategy & Career Course Lead
Golem Lab · LATAM
Reading Group Facilitator
Buenos Aires AI Safety Hub · Spanish-language seminar on constitutional AI
Ad Honorem Teaching Assistant
Iona University · Logic; Philosophy of Psychology & Neuroscience

Courses & Certifications

Honors & Awards

Stanford AI Learning Differences Hackathon — 1st Place
Stanford University
Cosmos Ventures Grantee
Cosmos Institute
Project ↗
University Innovation Fellow
Stanford d.school
Directory ↗
Dean’s Honors Scholarship
Iona University

Curriculum Vitae

I work at the intersection of ethics, political theory, and artificial intelligence — studying virtue, wellbeing, and agency in AI systems, and building the interpretability tools to study them. As a Philosopher and Computer Scientist, I examine how algorithmic architectures and institutional norms co-determine the moral trajectories of AI. My long-term goal is to help build the conceptual and technical foundations for an ethically responsible AI ecosystem — one that aligns machine capability with human flourishing.

Download PDF

Research Areas

AOSPhilosophy of Artificial Intelligence · Ethics of Technology · Philosophy of Mind

AOIEpistemology · Causation · Identity & Persons · Phenomenology & Consciousness · Plato

Education

PhD in Philosophy
University of Rochester
Virtue, wellbeing, and agency in AI systems, and the interpretability tools to study them.
BA in Philosophy & Computer Science
Iona University · Honors Program
Cybersecurity concentration · Magna Cum Laude.

Experience

CAMBRIA Fellow
Selective 3-week ML research program for AI safety: mechanistic interpretability and RL, based on the ARENA curriculum.
Founder & Board Member
SALVé–Techlanna · New York
AI enablement for enterprises and government — including building systems for Tucumán’s legislature in Argentina.
Digital Consultant
Hopps · Freelance
Advised and built for clients on the Webflow platform — implementation, optimization, and training on best practices.
Webflow Developer
QRoom LLC · Contract, Remote
Built and maintained The Futur’s site and a custom dashboard for Pro members; produced brand assets in Figma and Adobe CC.
University Innovation Fellow
Stanford d.school (Hasso Plattner Institute of Design)
Selected for Stanford’s design-thinking fellowship; prototyped a campus-wide collaboration app.
Resident Assistant
Iona University · New Rochelle, NY
Residential life: enforced housing policy, counseled peers, ran educational programming, and mediated conflicts.
Fellow
Hynes Institute for Entrepreneurship & Innovation, Iona University
Ran digital campaigns, served as an institute ambassador, and led a campus 3D-printing initiative.
IT Specialist
Quadro Diseño y Comunicación · Buenos Aires, Argentina
Led an award-winning team on digital content campaigns reaching 6M+ impressions nationwide.

Publications

Media

La ilusión del razonamiento artificial
Op-ed · Clarín · El Litoral (Argentina)
On Large Reasoning Models and Apple’s “The Illusion of Thinking” — why a model’s displayed step-by-step “reasoning” isn’t evidence of actual reasoning, and language manipulation shouldn’t be taken to imply thought.
Filosof.IA
Blog & podcast · Substack
Documenting my learnings across philosophy and AI.

Selected Projects

A regional AI safety community started in Rochester, connecting students, researchers, and builders across Central and Western New York. Launched with a $7,000 BlueDot Impact Rapid Grant (June 2026); its first event is a launch retreat for sixteen people, Nov 13–15 in the Finger Lakes, open to newcomers as well as researchers. The longer-term goal is a corridor-style network of connected AI safety hubs across Rochester, Syracuse, Buffalo, and Ithaca/Cornell.
An interactive map of the institutions, actors, and mechanisms that govern frontier AI — organized by governance layer and stacked by how much real enforcement power each wields. Built as a thinking tool for BlueDot’s Frontier AI Governance course (June 2026).
Activation-based classifiers that measure and steer moral dispositions in language models.
Open-source framework for detecting algorithmic bias in LLMs using established audit-study methodologies.
Cosmos Institute grant project: a framework for continually learning virtue-theoretic signals in multi-agent systems.

Talks & Presentations

Philosophical Problems of Algorithmic Agency
Escuela de Posgrado Newman — Perú · Slides (español)
Creativity as Interplay
Rutgers–Columbia Undergraduate Philosophy Conference · Paper (PDF)
The Why? Machine: Computing Causation
Iona University — Honors Thesis Day & Scholars Day · Video

Teaching

Graduate Teaching Assistant
University of Rochester · Philosophy of AI; Ethics of Technology
Guest Lecturer
Ethics of Technology (PHIL 120), University of Rochester · with Randall Curren · “Autonomous Weapon Systems”
Guest Lecturer
Philosophy of Wellbeing (PHIL 239), University of Rochester · “Work and Play in the Shadow of AI”
AGI Strategy & Career Course Lead
Golem Lab · LATAM
Ad Honorem Teaching Assistant
Iona University · Logic; Philosophy of Psychology & Neuroscience

Certifications

Frontier AI Governance
BlueDot Impact · Credential
CS50AI: Introduction to Artificial Intelligence with Python
edX · HarvardX · Certificate · Credly

Honors & Awards

BlueDot Impact Rapid Grant — CNY AI Safety
BlueDot Impact · cnyaisafety.org
Stanford AI Learning Differences Hackathon — 1st Place
Stanford University
Cosmos Ventures Grantee
Cosmos Institute · Project
University Innovation Fellow
Stanford d.school · Directory
Dean’s Honors Scholarship
Iona University

Graduate Coursework

Theory of Knowledge
PHIL 503 · with Earl Conee
Topics in Philosophy of Mind
PHIL 544 · with Grace Helton
Selected Topics in Philosophy of AI
PHIL 557 · with Rush Stewart
Selected Topics in Ancient Philosophy
PHIL 465 · with Randall Curren
Data, Algorithms, Justice
PHIL 435 · with Mark Povich
Selected Topics in Modern Philosophy
PHIL 470 · with Dante Dauksz
Selected Topics in Social and Political Philosophy
PHIL 523 · with Randall Curren
Seminar on Artificial General Intelligence
CSC 409 · audited · with Christopher Kanan
Metaphysics
PHIL 442 · with Paul Audi
Philosophy of Artificial Intelligence
PHIL 457 · with Jens Kipper