Inverse reinforcement learning promises to tell us how agents will behave in worlds they have never seen. The trouble is that no single agent has seen enough of this one. The fix, it turns out, is to stop studying them one at a time.
Suppose a city shuts down a train line. Every commuter who relied on it will re-route, but not in the same way. One cares mostly about speed and will pay for a car. Another will take a slower bus to save money. A third will do almost anything to avoid a packed platform. A planner who wants to know what happens next has only a record of what people did while the line was running. That record is behavior, and behavior is exactly the wrong thing to extrapolate. It is the joint product of what people want and the world they were wanting it in, and the world is precisely what is about to change.
Economists know this trap by name. In 1976, Robert Lucas argued that the statistical regularities in historical data would break the moment policy tried to lean on them, because people’s decision rules are themselves optimized against the policy environment. Forecast with decision rules and you will be wrong; forecast with the preferences and constraints that generated them and you have a chance. Inverse reinforcement learning is the same idea in different clothes. Behavioral cloning learns the policy. IRL learns the reward that rationalizes it—the agent’s revealed preferences—and the reward, unlike the policy, should survive a change in dynamics. Hand the recovered reward a new transition kernel, re-plan, and you have a counterfactual.
That is the promise. The problem, which anyone who has tried this will recognize, is coverage. A rational agent spends its time where its reward is, so its demonstrations pile up in a narrow sliver of the state space. A reward inferred from that sliver is pinned down where the agent went and nearly arbitrary everywhere else. And a changed environment, almost by definition, sends agents everywhere else: close the train line, and commuters flood routes they rarely touched. The states you most need to evaluate are the ones you understand least.
In the maximum-entropy framework, the tension even comes with a dial. Turn the temperature up and agents wander widely, but their choices flatten toward uniform and say little about what they prefer. Turn it down and choices become sharp and informative, but the agents stop exploring. Fitting N agents independently doesn’t escape this; it inherits the problem N times.
···
But a planner never watches one commuter. They watch thousands, all moving through the same network, all wanting overlapping things—speed, cost, comfort—in different proportions. Our paper takes that overlap seriously. If each agent’s behavior is a different mixture of a small number of shared ingredients, then an agent who never visited a state can borrow what others revealed there. Each agent need not visit every state, as long as somebody does.
The interesting decision is where to put the structure. The instinct is to assume that rewards are low-rank. We put it on the policies instead—specifically on the matrix of anchored action logits, one row per agent and one column per state-action pair, factored as Θ = WΦ: k shared basis functions, and a vector of k loadings for each agent. (This is a structural assumption, not a consequence of linear rewards; the max-ent optimality map is nonlinear and can raise the rank.) Three things follow from that choice, and they are why this is more than a regularizer bolted onto IRL.
First, the assumption is about something you can see. Choice probabilities can be estimated from data; rewards cannot. So the rank is testable, and we test what happens when it’s wrong: recovery degrades gracefully when k is too small and plateaus when it is too large, rather than falling off a cliff.
Second, it turns the learning problem into a familiar one. Fitting Θ is a low-rank multinomial logit—matrix completion with a softmax likelihood and a nuclear-norm penalty—which comes with a mature statistical toolkit.
Third, and most consequentially, it makes planning cheap. Under maximum entropy, and a normalization that resolves the usual ambiguity about which reward produced a given policy, the reward is an affine function of the log-policy: ri = Lμ(log πi − g) + g. The operator Lμ hides a Bellman fixed point—the expensive step in IRL, ordinarily re-solved for every agent. But if the logits are linear combinations of k basis functions, you can push each basis function through the operator once and recover every agent’s reward as the same linear combination of the results. That is k fixed points instead of N. A new agent costs k numbers, not another planning problem. (The exact log-policy includes a log-sum-exp normalizer that is not low-rank, so the implementation applies the operator to the raw logits. We characterize the gap exactly and measure it directly; on accuracy it is a wash.)
···
What does it actually take for pooling to work? Here the theory says something sharper, and more interesting, than “more data helps.” The coverage condition is not a visit count. At each state, what matters is whether the agents who pass through it collectively span the latent directions of preference. If everyone who visits a state has similar tastes, the data there cannot distinguish the ways in which the other agents differ. Errors along those unspanned directions never register in the training loss. They hide, and then they surface off-support.
Formally, the requirement is a restricted-eigenvalue condition on the source design over the cone of possible low-rank errors: the demonstrations must span the low-rank cone. One worked example has an agent–state pair that is never observed at all, yet whose reward is fully recovered as data grows, with target regret vanishing even when the target policy sends the agent straight into that unseen state. The finite-sample bound scales like the square root of k/n, up to logarithmic factors—rank, not agents times states times actions—divided by that coverage constant, and it carries the estimation error all the way through the reward operator to regret, and to occupancy-weighted KL against the true soft-optimal policy in the new environment.
Then we did something unusual for a machine-learning paper: we ran a causal experiment on our own estimator. At selected states, we progressively deleted observations from exactly the agents that contribute most to the state’s leading eigendirection, alongside volume-matched controls that lose the same number of observations without regard to alignment, and placebo states that lose nothing. Error along the targeted direction grows monotonically with the dose, and a difference-in-differences puts the damage in that direction and not the others. It isn’t how much data you have. It’s whose.
···
The experiments use three synthetic worlds, chosen so that ground truth is known: a four-room gridworld where agents collect colored objects under partial observability; a highway simulator whose drivers trade off speed, collisions, headway, and lane changes (some of them, it must be said, like to tailgate); and a recommender in which users burn out on topics. No method ever sees the reward features.
The telling measure is off-support recovery: how well each agent’s reward is recovered at states that only other agents visited. Across budgets from a hundred to ten thousand decisions per agent, the low-rank method leads in all three domains, by the widest margin when data is scarce. Per-task methods catch up only on the highway at the largest budget, where enough within-agent data eventually supplies its own coverage. Pooling everyone into a single reward fails for the obvious reason—it buys coverage by throwing away heterogeneity—and clustering agents into a handful of discrete intents proves too coarse. More striking: outside their own support, single-task estimates can get the direction of reward variation backward. Pooling gets it right.
The transfer experiment speaks most directly to the train line. In the recommender, where exact planning allows comparison with an oracle, we increasingly suppress each user’s favorite genre, forcing behavior into states that user rarely visited. Behavioral cloning and IQ-Learn hold their own when nothing changes; then cloning falls away as the target environment drifts from the source. As occupancy shifts off-support, IQ-Learn and GenPQR lose roughly seventeen points of normalized return for every tenth of occupancy mass that moves. The low-rank method’s return stays flat, with a slope statistically indistinguishable from zero.
And it scales. Growing the population from eight agents to a hundred and twenty-eight at fixed rank, per-task baselines slow down roughly linearly; the low-rank method’s wall-clock grows as N0.21. At 128 agents it runs four to thirteen times faster than per-task GenPQR, which uses the same reward operator, while recovering off-support rewards with correlations of 0.88, 0.53, and 0.78 against GenPQR’s 0.74, 0.39, and 0.18. A basis learned once also travels: learning it from 128 agents rather than eight raises off-support recovery for newly arriving agents by about 0.15 in correlation.
···
There are honest limits. The guarantees cover tabular MDPs with a known source kernel, while the experiments use neural function approximation; extending the theory to that approximation error, and to the normalizer we approximate away, is open, as is relaxing the shared reward normalization. The environments are synthetic—necessarily, since you cannot grade reward recovery without ground truth—so validation on real behavioral data remains to be done.
What we find most appealing is the reframing. The usual picture of IRL is a solitary inference: one expert, one set of demonstrations, one reward, and a long struggle against the edges of the data. Here, a population becomes an instrument. Each agent explores the corners of the world its preferences lead it to, and the differences among them—which look at first like a nuisance, N problems instead of one—turn out to be the asset. The commuter who has never once taken the crosstown bus has, through everyone who did, told you something about how they would ride it.