Finding Superstars
Author: James Brand
Date: 2026-06-29
I’ve been interested lately in how firms identify top/superstar talent in technical skills, especially now given the big AI talent conversations going on. A recent post by Epoch AI and this comment by Roon on Twitter rekindled my interest.
The Epoch post makes references to a classic paper on the economics of superstars and gives examples like track stars and artists (Taylor Swift) as markets in which the very top performers massively outperform even second place. I’m not convinced these are particularly useful for what we see in the AI/technical talent space. Taylor Swift makes more than her peers because the products she makes are popular; AI superstars are paid amazing sums because their firms hope they will make something popular. That uncertainty is a key feature of labor markets, and even more important the rarer a signal a firm is hunting for – firms make consequential hiring choices based only on a small number of signals, and when they are looking for a top-10 candidate world-wide, my intuition is that they are going to miss the mark a lot.
So, I find myself both believing that the very best technical talent is worth 10x the median (and probably 5x the 95th percentile), and that finding that talent within current hiring structures must be basically impossible. Even outside of the tails of the distribution, I think everyone knows examples of people who are severely over- or under-valued/paid, and surely this must get worse in the tails mechanically. And yet, we see the AI labs spending a ton for their top talent – and I like to believe they must be doing this somewhat rationally – so, this post is me thinking about how hard it’d be to rationalize it.
Simple model
Starting with the simplest model I could think of, imagine a world in which every person has some average level of “idea quality” which is distributed normally over the population, and that each idea a person has is normally distributed around their mean (visualized below). The goal is to find the extreme upper end of intelligence/ideas, in the form of “superstar” thinkers.
The key determinant of how easy it is to find these superstars is the ratio of within-person quality dispersion to dispersion between people. I have no idea what that should be here, but in my experience even the smartest people often have some bad ideas, so I would think that ratio is only a little less than 1. For reference below, a value even as low as 0.5 would basically mean no superstar ever has an idea as bad as the median person’s best idea, which I think would be very generous to the former.
My mental model of hiring here is essentially: you get to see some number of “ideas” from a candidate, and from that you have to decide if they meet your hiring bar. Let’s imagine, generously, that businesses know the full quality distribution (so that when they say they want a “top-1%” candidate, they know what that means quantitatively). Below, I calculate false discovery rates for naive screening rules based on (1) the average observed idea, and (2) the best observed idea, which I think is more common at the very far right tail of individual talent. In these simple rules, firms say “I will take any candidate whose observed skill exceeds the level I’m targeting.”
What these plots show, albeit somewhat densely, is that if you either (a) only see a few ideas for each candidate, or (b) only observe the best idea out of a large set and screen on that naively, you’ll have consistently bad false discovery rates that become horrendous as you move into superstar territory. Even if you get to see 100 ideas play out, and even if within-person quality is distributed pretty tightly, you might mis-flag a top-0.001% superstar 10% of the time. Given the size of compensation packages reported (worth tens or hundreds of millions of dollars), that seems like a risky bet to me. Worse, my read is that people are often known for and judged by the most successful idea with which they are affiliated, so I think the bottom row of graphs is closer to reality than the top. 1
Another way to look at this is to ask: if I evaluate two top-1% candidates, which are much easier to find, how often will I rank them incorrectly in this model? If we assume in-person dispersion is half of cross-person then the answer (shown below) is that if a single idea is seen by the firm, they do a little better than a coin flip at ranking any two candidates. Even after 100 observed ideas, 10% of pairs get mis-ranked. Again, that seems like a scary proposition if we were trying to hire an employee worth $50M/year.
You can also see how mismeasurement and luck in this model could exacerbate the mismatch between compensation, or perceived skill, and actual skill. You start your career and get a lucky win, which becomes the only (publicly) visible idea you’ve had and someone hires you at higher compensation. Being hired into that role serves as evidence of your skill for the next role, and so on. This is all compounded by the fact that validating the quality of ideas is almost as difficult as coming up with good ones in the first place, and by the many organizational frictions that prevent good ideas from trickling through organizations.
Rationalizing
So obviously this isn’t how the world works, but hiring firms do rarely get to see a huge portfolio of (solo) projects/ideas/output from candidates. This seems really important when we look at the market for AI researchers who are increasingly unable to publish findings, meaning hiring firms have little public information to gauge skill, but also true of many technical skillsets. So, how would we have to tweak our world model to make sense of the superstar hiring that is actually happening?
Heteroskedasticity
One option is that the variance or shape of idea quality might become more favorable for high-skill thinkers. Maybe they essentially stop having bad ideas (left) or have generally lower quality variance (right)?
Maybe this is true directionally, but I’ll believe it when I meet the genius that never has bad ideas.
Winner-take-all
An explanation I hear lately is that the stakes in AI markets are winner-take-all so that superstar hiring makes sense, even if you are wrong about all but one of the superstars you try to hire, because the single one you correctly identify lets you win the market. You can imagine, in the extreme, hiring all possible superstars to ensure you win the market. I am, to say the least, skeptical of this argument in equilibrium, especially given the many online posts I’ve seen about how leaky the big labs are. Hard to argue for a winner-take-all secret sauce in that kind of world. Big labs also compete for perceived superstars, which shifts much of the potential rents to the labor side and makes the net value of the superstar even more tenuous.
The applications of this idea that I’m much more optimistic about are the competitions that firms like OpenAI (“parameter golf”) and Jane Street have run, which aim to find hidden talent via wide searches rather than looking at existing top researchers. Spending a fraction of superstar compensation on hiring a few dozen such researchers could easily be higher value than a single superstar simply from first-order statistics – if you think you need the best possible draw from a distribution to win, take as many as you can!
Misc. other options
- Maybe technical superstars have massive spillovers? Take for example Andrej Karpathy, who was hired by Anthropic and is known in part for instructive YouTube videos and toy git repos that teach people about AI. Maybe he makes everyone on the team 5% better and that value compounds?
- Do scientists go on hot
streaks? Maybe there are temporal windows in which superstars will
go on a run of amazing ideas and hiring during those windows helps
increase signal-to-noise ratios– on the other hand, the same paper I
linked could be interpreted pessimistically to say it’s even harder than
you’d expect to know when a superstar will succeed. Sometimes you’ll
hire someone after 5 amazing ideas and their next 10 will be duds!
- Extra cash and expensive capital – many have noted the amazing sums
being spent by clouds and AI labs on compute lately. If the labs are
sitting on cash but face hard constraints in the amount of capital they
can buy, maybe superstar hiring comes from a simple substitution story.
Not sure this is enough, but I could see it as part of the story.
- Network effects: I do worry that some of what happens in technical/academic groups is a sort of circular logic of skill evaluation. People like to hire people they’ve worked with before, or heard of through papers, and a lot of that is random. I haven’t seen perfect evidence of the ways randomness in academics flows through to citations and prestige, but anecdotally it’s a big effect and there are a couple of papers touching on it arguing that early coauthorship with other superstars and affiliation with prestigious institutions dramatically shape citation outcomes, which I suspect is the result of both real skill differences and bad citation dynamics and incentives in publishing.
End
The last substantive thing I’ll mention is that my previous post asked whether AI reduces the costs of skill verification. While I still like that idea, it’s way less likely to work for superstars generating new knowledge for the same reason that it’s hard to build AI versions of those superstars. The value long-term of new knowledge, new ideas, or new products, especially in a market setting, is difficult to verify even ex-post given the many complexities involved.
I’m not convinced that any of these are sufficient to explain the compensation packages that have been reported, but I suspect that’s because my beliefs about the eventual market size and winnable profits in AI differ from those of the people doing the hiring. Beliefs are critical and in a world of giant and growing valuations, it’s easy for reasonable people to disagree about the size of the pie. If I believed (I do not) that the eventual market equilibrium would entail a single firm capturing most value, or that the eventual costs of AI inference would decline sufficiently to generate giant profits (unsure), then I might think taking a few big labor market bets makes sense. Of course, if any labs are interested in hiring a $50M economist, please let me know!
The economist in me wants to note that fully rational firms could in principle adjust their screening rules to reduce false discovery. E.g., a firm that knows the distribution of within- and across-person quality and an estimate of the number of measured ideas over which they are observing the best one could raise their threshold to account for this sampling process. In practice, that seems extremely unlikely, because to shrink the false-positive rate sufficiently would require reviewing a huge number of candidates intensively.↩︎




