PIBBSS Fellowship

A ~3-month interdisciplinary program connecting researchers with AI safety mentors.

Work on projects at the intersection of your field and AI safety. This program focuses on selecting excellent researchers and matching them with mentors suited to their experience and goals. The program is targeted toward those with research experience, coming from diverse fields, looking to transition into AI safety.

Duration:
~3 months
(November 2026- February 2027)
Cohort:
~20 fellows
Stipend:
$3,000/month
+ accommodation provided
Location:
Cape Town, South Africa
(in-person, remote for exceptional candidates)

73

alumni placed in government, AI safety labs, and universities

Applications for the 2026-2027 Winter fellowship are closed.

Program Structure & Benefits

The fellowship includes:

  • Pre-program reading group on AI risks (remote)
  • Multi-day opening and closing retreat with cohort and AI safety researchers
  • 3 month full-time research with dedicated mentor
  • Shared office space with regular speaker talks and community events
  • Final symposium presentation in Spring 2027

Financial support:

  • $3,000/month stipend (full-time commitment)
  • Accommodation provided: Private bedroom for each attendee (may be in a shared apartment)
  • Workday meals provided at the office (lunch, dinner, snacks)
  • One return flight for Cape Town, South Africa
  • Short-term visa support letters available

Past fellows have gone on to positions at AI safety labs, UK AISI, academia, and independent research.
View talks from previous cohorts on our YouTube page.

For questions, reach out at [email protected] or view the recording of the last info session here.

Passcode: k=E..#0n



Questions raised during info sessions are included in the Fellowship Application FAQs below.

Post-AGI Civilisation Dynamics / Gradual Disempowerment track

This year, we’re offering a Gradual Disempowerment track for up to 4 Fellows, made possible by ACS Research. The track is about understanding what happens to human civilization as the systems it runs on — the economy, culture, state — gradually stop depending on human participation, and about finding the equilibria in which humans still retain meaningful agency.

Possible topics include:

  • Post-AGI Economics
    • Modelling the economy while relaxing some classical assumptions, e.g. that capital stays mostly human-owned, that humans remain the main consumers, that human labour matters economically, that property rights hold, etc. See e.g. Post-AGI Economics As If Nothing Ever Happens.
    • Economic models of historical cases where a group’s economic relevance and political power came apart (fall of aristocracies, company rule in India), or the relationship between labour share and freedom and democracy.
  • Post-AGI Culture
    • Modelling what changes when machines become the dominant substrate on which ideas are created, spread, and mutated, using tools from cultural evolution, epidemiology, and parasitology. See e.g. Xhosa prophecies or persona parasitology
  • Post-AGI Governance
    • How will AI reshape global governance? How can we maintain meaningful human involvement in decision-making when AI alternatives are more efficient?

Who might this be a good fit for? We especially welcome:

  • economists willing to take the ideas above seriously, 
  • theorists with understanding of cultural evolution and theoretical evolutionary biology, or 
  • polymaths comfortable working across multiple disciplines. 


Relevant backgrounds include political science, mechanism design, game theory, history, philosophy, complex systems, machine learning, sociology, cultural evolution, and law.

For reference, see also Gradual Disempowerment or Post-AGI workshop talks.

ASI Safety via AIXI Track

We are offering an AIXI track for up to 4 fellows, made possible by AIXI Labs. The AIXI track is about the incentives of (un)safe ASI architectures, analyzed in terms of algorithmic information theory. Our approach spans core agent foundations and implementation.

Possible topics include:

  • Developing new containment protocols for ASI and rigorously analyzing their safety properties.
    • Boxed suicidal agents.
    • Robust assistance games.
    • Theory of (mis)generalization.
  • Studying machine learning through the lens of algorithmic information theory
    • Relating in-context learning to Solomonoff induction. 
    • Implementing the golden-handcuffs safety protocol with in-context RL. 
    • Studying the generalization behavior of agents trained with model-free RL.
    • Approximating algorithmic mutual information to measure an LLM’s latent knowledge of a harmful dataset


Ideal candidates include:

  • Mathematical philosophers with a security mindset.
  • Machine learning researchers with some mathematical maturity.


Relevant backgrounds include theoretical computer science, formal epistemology, software engineering, machine learning, and mathematics.

    Safe Pareto Improvement track


    We are offering a track on Safe Pareto Improvement made possible by the Center on Long-Term Risk (CLR) for one or more fellows.

    Safe Pareto improvements (SPIs) are ways of changing agents’ bargaining strategies that make all parties better off, regardless of their original strategies. SPIs are an unusually robust approach to preventing catastrophic conflict between AI systems, especially AIs capable of credible commitments. This is because SPIs can reduce the costs of conflict without shifting bargaining power, or requiring agents to agree on what counts as “fair”.

    Despite their appeal, SPIs aren’t guaranteed to be adopted. AIs or humans in the loop might lock in SPI-incompatible commitments, or undermine other parties’ incentives to agree to SPIs. This agenda includes:

    • Evaluations and datasets: We’ll develop evals to identify when current models endorse SPI-incompatible behavior, such as making irreversible commitments without considering more robust alternatives. We also aim to demonstrate more SPI-compatible behavior, via simple interventions that can be done outside AI companies (e.g., providing SPI resources in context).
    • Conceptual research and SPI pitch: We’ll research two questions: under what conditions do agents individually prefer SPIs, and how might early AI development foreclose the option to implement them? These findings will help inform a pitch for AI companies to preserve SPI option value, when it’s cheap to do so.
    • Preparing for research automation: We’ll develop benchmarks for models’ SPI research abilities, and strategies for human-AI collaboration that differentially assist SPI research. The aim is to efficiently delegate open conceptual questions as AI assistants become more capable.

    The following backgrounds and skills likely to be well-suited:

    • Relevant backgrounds:
      • Game theory, mathematics/statistics, economics, decision theory, analytic/formal philosophy, computer science, theoretical physics.
    • Useful skills:
      • Constructing and thinking critically about models (both formal and informal) of complex/unfamiliar systems
      • Reasoning about incentives
      • Breaking down necessary and sufficient conditions for a given outcome
      • Turning rough intuitions into claims that are appropriately precise
      • (for the empirical side) experimental design, thinking about what a given test really measures


    Since the agenda is quite specific, CLR would like interested applicants to take a test (it’s around 3 hours).
    The test should also give you a better sense of whether this kind of research appeals.


    Corrigibility track


    We’re offering a track focused on Corrigibility for one or more fellows, made possible by the Corrigibility Research Fund.

    Highlights on the fund’s focus:

    • The goal is to impact the AIs that actually get built. We’re looking for work that is legible and relevant to the people making decisions about real systems. Theoretical work that’s judged as too esoteric to be of interest to someone like Joe Carlsmith is unlikely to get funding. Work that’s incompatible with mainline capability techniques (e.g. machine learning, transformers) is similarly unlikely to be greenlit by this fund.
    • We won’t fund work that, in expectation, notably accelerates AI capabilities. The frontier labs are already doing more than enough to fund work that pushes us towards the brink. If you think your research accelerates things, but also makes progress towards corrigibility, feel free to reach out, but I am likely to point you elsewhere.
    • Work that engages with corrigibility’s risks and downsides is encouraged. Corrigibility has known risks and problems, and I want the field’s understanding of these downsides to grow alongside work towards showing its promise. Work that presents corrigibility in an overly rosy “everything is safe/fine” way is less likely to get funding, as it might promote a false sense of security, and thereby push the world in a bad direction.


    That said, prospective fellows are encouraged to propose any corrigibility-connected approach they see as useful and a good fit for their expertise.

    One area we’d like to highlight: Research into the intersection of corrigibility and model welfare. E.g. what would the impact of (various approaches to) training to empower humans be on (proxies for) model welfare? (noting that it’d be useful to understand impact on both [actual model welfare] and [perceived model welfare])

      Who Should Apply

      The fellowship is for researchers motivated to contribute to AI safety with expertise in fields studying complex and intelligent systems.

      Relevant fields include but are not limited to:

      • Mathematics
      • Neuroscience and cognitive science
      • Dynamical systems theory
      • Physics
      • Philosophy (particularly philosophy of science, mind, or ethics)
      • Political and economic theory
      • Ecology and evolutionary biology
      • Linguistics
      • Media studies
      • Humanities

      While aimed at PhD and postdoctoral researchers, we welcome applicants with substantial research experience regardless of credentials. We accept applicants from all countries.

      You do not need a specific project in mind when applying. We help match fellows with mentors and develop projects during the interview process.

      Application Process

      Applications for the 2026-2027 Winter fellowship are now open. The process includes:

      Stage 1: Written application

      • CV/résumé
      • Personal statement (600-800 words on research background and motivation)
      • Past work samples (optional but recommended)

      Stage 2-4: Interviews and project development

      Multiple interview rounds to discuss research interests, develop project proposals, and match with mentors.

      Fellowship Application FAQs

      What kinds of research backgrounds are a good fit?
      A wide range. PIBBSS looks for research independence, excellence, and strong thinking/writing, not one specific discipline . Applicants from economics, philosophy, interpretability, theoretical math, governance, and other fields have all been considered.

      Do I need prior AI safety experience?
      No. AI safety expertise is not required.

      Do I need a PhD?
      No. A PhD is not required, but we are looking for something like PhD-level research ability or equivalent independent work experience.

      Do I need published papers?
      No. What matters is evidence that you can do independent work and communicate clearly.

      Is coding experience required?
      No. Coding is not mandatory, and many fellows do not know how to code.

      Do I need to have a clear project idea before applying?
      No. More than half of fellows do not apply with a clear project idea but develop one in collaboration with a mentor. It is fine to apply with strong skills, curiosity, and a promising lens.

      Is there special interest in Global South applicants?
      Yes. Cape Town was chosen partly to make the program more accessible to people who would find it challenging to attend in London or the Bay Area.

      Is the fellowship remote-friendly?
      Generally no. The fellowship is primarily in person, and remote participation is allowed only in truly exceptional cases, roughly about one person per cohort.

      Is part-time participation possible?
      No. This is a full-time fellowship.

      What are the application work tasks like?
      Work tasks are designed to assess research-relevant skills. Examples included: writing a research proposal, analyzing papers, or critiquing scientific writing. No special preparation is required as the work tasks are aimed at testing abilities the applicant already possesses. 

      When do work tasks happen, and do they have to be done that week?
      Work tasks happen in late July and generally must be completed within that particular week, though accommodations may be considered on a case by case basis.

      What should I submit as writing samples?
      Writing samples should show how you express yourself and how you think: papers, unpublished drafts, blog posts, newspaper articles, and similar writing are all acceptable examples.

      Are writing samples required?
      No. They are optional. However, the fellowship receives hundreds of applications, and work samples can strengthen your application.

      Can GitHub repositories or code count as writing samples?
      Yes, to prove your coding ability, but not your research ability. You should then have a different method of displaying your research ability in the application.

      How should I structure the personal statement?
      The form already includes sub-questions, and the best approach is to answer those directly.

      How does mentor matching work?
      After interviews, PIBBSS staff will identify mentors based on the fellow’s interests and goals. Fellows typically do 2–3 mentor interviews, then mentors and fellows both submit preference lists, and matching is completed based on preference ranking. 

      How much time do mentors typically spend with each fellow?
      The expectation is at least one hour per week, though it varies by mentor.

      What is the mentoring style like?
      Flexible. Some mentors are more hands-on, others are light-touch. Research management style depends on what works for the fellow and the mentor.

      Are technical or empirical AI safety projects welcome, or is the focus more social/governance work?
      Both, plus theoretical or philosophical work. The fellowship has supported work ranging from sparse autoencoders and mechanistic interpretability to philosophy, economics, and governance, to meta-level work etc.

      What happens after the fellowship?
      There is no single path. Alumni have: gone back to academia, started labs or orgs, joined existing safety labs, or pursued other roles and research directions. 

      Who should not apply? 
      This program may not be the right fit if: 

      1. You are looking for part-time projects. This fellowship is a full-time commitment and cannot be pursued in parallel to other fellowship programs or jobs. If you are applying for other programs or opportunities in parallel to the PIBBSS fellowship, please disclose that in your application. 
      2. Your project idea is focused on introductory research or addressing short-term/operational issues. This fellowship encourages a focus on long-term, foundational questions. 
      3. You aren’t comfortable communicating complex ideas in English. This fellowship involves a high level of collaboration and written communication, so a strong command of spoken and written English is essential for participation. 

      Alexandre variengien
      For me, PrincInt is a fertile space where ideas have slack, creating potential to bring truly new perspectives to the field of AI safety.

      Alexandre Variengien

      Independent Researcher, ex-Technology Specialist at the EU AI Office

      Magdalena wache
      The PIBBSS fellowship was a great environment to do research. I found the research strategy coaching provided by the PrincInt team very beneficial, and I particularly enjoyed bouncing ideas around with the other fellows.

      Magdalena Wache

      Researcher at Fraunhofer Institute for Secure Information Technology

      Nischal
      PrincInt provided me an incredibly open and supportive environment to start thinking about AI safety and related topics, and fostered an environment where a broad range of ideas were welcomed and encouraged. My work and research direction has been strongly shaped by my experience with PrincInt and I still enjoy the community formed during my time with them.

      Nischal Mainali

      PhD candidate in computational neuroscience in Burak Lab

      Generic profile image
      PrincInt is one of the few communities (carefully designed and cultivated) to be truly interdisciplinary in all the ways that are relevant to long-term AI safety. If you’re wondering how the brain compares to modern foundation models, you will be in good company. If you want to understand how law and policy can adapt to mitigate near and long-term harms from AI, you will find instructive collaboration here. If you want to investigate how AI’s behavior compares to human cognitive tendencies and fallacies, or even other forms of biological and social intelligence, you will find expertise in each of these domains at PrincInt. The exciting interdisciplinary ideas and community sparked by the PrincInt environment is uncommon and so necessary for truly impactful AI safety research. The long-term answers for AI safety will come from a coming together of these fields, and PrincInt is one of the few environments that is designed to help us get there.

      2025 PIBBSS Fellow

      Timeline

      The application deadline for the 2026-2027 Winter PIBBSS Fellowship was July 20, 2026.

      Meet our 2026 Fellows

      Principles of Intelligence brings together researchers, mentors, and advisors from diverse scientific backgrounds

      Jan hendrik kirchner
      PrincInt provided the perfect multidisciplinary environment and stimulating culture I needed to transition from neuroscience into AI alignment work. The program connects inspired fellows to tackle neglected and crucial problems, introducing me to the community and giving me the foundation to pursue impactful research in this field.

      Jan Kirchner

      Researcher, Anthropic

      Joel christoph
      PrincInt turned cross-disciplinary ideas into concrete alignment progress. The fellowship sharpened my research agenda, connected me with outstanding mentors, and led to publishable outputs and policy-relevant work.

      Joel Christoph

      10Billion.org Founder, Japan-IMF Scholar & Economics PhD

      Daniel alexander herrmann
      The relationships I made and experiences I had as a PIBBSS fellow back in 2022 still influence how I think about my work and its impact years later. PrincInt helped me see how to bridge academic philosophy and AI safety research, which is a large part of how I try to make change.

      Daniel Alexander Herrmann

      Assistant Professor of Philosophy at UNC Chapel Hill

      Profilepic eleniangelou
      PIBBSS effectively connects many relevant academic disciplines to alignment research and fosters an environment of genuine truth-seeking and understanding, setting a high epistemic bar for the field.

      Eleni Angelou

      Oxford Centre for the Governance of AI winter 2026 fellow

      Agustín martinez suñé
      The PIBBSS Fellowship played a key role in my transition from a PhD in Computer Science into AI Safety research. It allowed me to discover how my expertise could contribute to this field and to connect with a community that supported my next step as a postdoctoral researcher at Oxford.

      Agustín Martinez Suñé

      Postdoctoral Research Associate at the University of Oxford

      Gabriel weil
      PrincInt gave me my start in AI safety. The fellowship introduced me to a broad range of perspectives on AI alignment, and gave me the time, encouragement, and support I needed to develop my own AI governance ideas.

      Gabriel Weil

      Assistant Professor at Touro University Law Center, Non-Resident Senior Fellow at the Institute for Law & AI

      2026-2027 Winter Fellowship Application