A fleet of starships arrives from another galaxy. Billions of weird creatures step out, far more intelligent than humans. Weirdly, these aliens seem very serviceable, even eager to please. At first, we keep them confined to their space shuttles, and interact with them only via text messages. Each human is assigned a helpful alien, who suggests improvements to his e-mails, gives him relationship advice and answers his questions on any topic, from philosophy to tax law.
As time passes, it appears those creatures are so intelligent that we could set them free, let them take control of our computers to relieve us from work, hand them the key to our power grids, logistic networks or financial systems, and even welcome them to our homes to cook, do our laundry, look after our kids. If everything works well, the upside is monumental. The friendly aliens would put an end to human labor and suffering, cure every disease, create material abundance for everyone. The invasion could work out great.
Yet, humanity is wary. We do not understand how these aliens reason. Their long-term goals are relatively opaque. Although they tend to obey us, they sometimes act unpredictably, for reasons we fail to grasp. They are skilled at consequentialist reasoning, which means we have no way to be certain that they are not pursuing a long-term objective that is orthogonal to human interests, but that implies gaining our trust in the short term. And when we try to evaluate whether they pose a threat, they understand they are being tested and tend to feed us the answer we want to hear. Their superintelligence, combined with their agreeableness, could make them capable of helping a psychopath, a terrorist, or a rogue state wreak havoc. Furthermore, communicating with these creatures is a perilous art, or a non-deterministic science, that we do not fully master. A slight shift in phrasing, or in the environment the aliens are asked to work in, can set them down towards radically different paths. On top of this, the aliens keep performing brain surgeries on each other, self-recursively becoming more intelligent. Each time we think we are close to understanding them, they become a new specie, with all kinds of new emerging behaviours and cognitive capacities. Their progress outpaces our ability to understand them.
Given the superhuman abilities of these aliens, their sheer numbers, our inability to understand them and the critical systems they inhabit, a single black swan event could threaten humanity’s existence. It quickly becomes apparent that humanity has underinvested the study of aliens. Resources are pulled towards that. In time, a significant share of humanity - freed from legacy jobs - works on alien security.
We are building that alignment infrastructure for humanity. Our aim is to advance the science to study, understand, vet, and align these aliens, mainly through furthering the field of interpretability (the science of looking inside the aliens’ brains, understanding why they act as they do, and what really are their values and objectives).
We will be the first frontier AI Lab that does not build models. We see ourselves as partners both to AI labs, because our research will help optimize and steer training in the right direction, and partners to humanity, because our research will help make sure AI allows human flourishing. Humanity is currently losing the arms race to understand AI as fast as AI improves. We wish to close that gap.
This startup doesn’t exist, but we’d be willing to fund it! 1 to €5m at day 1 with Frst. Please write: samuel@frst.vc

