Skip to content

Meta Open-Sources Rebalancer, a Generic Assignment-Problem Solver

Meta open-sourced Rebalancer, the assignment-problem solver it has used internally for nine years to place racks, servers, and tasks across its datacenters.

Meta Rebalancer graphic: a balance-scale symbol with blue and green blocks connected by curved lines on white
Rebalancer · Credit: Meta

Meta open-sourced Rebalancer, an assignment-problem solver the company has run internally for more than nine years to place racks, servers, and tasks across its datacenter fleet. The release covers the core library under an Apache 2.0 license, alongside a companion debugging tool called Rebalancer Explorer.

Assignment problems ask a deceptively simple question: given a set of objects and a set of bins, how should the objects be distributed to satisfy constraints while optimizing a goal? Meta says it hits this pattern throughout its infrastructure stack: spreading racks across electrical fault domains, packing servers into services for fault tolerance, allocating tasks onto machines within resource limits, and routing user traffic to the nearest datacenter. According to Meta, most reusable optimization frameworks stumble on one of two problems: translating real policies into rigid mathematical formulas, or scaling past what commercial solvers can handle once the problem gets large.

Rebalancer's answer is to separate how a problem is described from how it gets solved. Engineers define objects, bins, constraints, and objectives through a layered API; Meta cites a task-placement example where servers sit inside rack-level "scopes," CPU and storage are modeled as "dimensions," and a spec enforces that each rack hosts only one job type. That specification compiles into a directed graph, which Rebalancer then feeds to either an exact solver (commercial tools like FICO Xpress and Gurobi, or the open-source HiGHS) for small and mid-sized problems, or a parallelized local-search heuristic for the largest ones. Meta says almost all of its large-scale allocation work runs on local search, while exact solving is reserved for smaller jobs or for tuning a local-search baseline offline.

The scale Meta describes is substantial: Rebalancer handles roughly 40 million assignment problems a day across more than 30 distinct problem formulations, spanning uses from shard-to-server placement (Shard Manager) and server-to-service allocation (RAS) to edge traffic routing (Taiji) and even meeting-room and desk assignments. Meta reports a P99 solve time of 12 seconds for a 265,000-object, 3,200-bin problem, and an average of 171 seconds across more than 3,400 runs on problems exceeding a million objects. The underlying research is documented in a paper Meta co-authored for OSDI 2024, "Optimizing Resource Allocation in Hyperscale Datacenters: Scalability, Usability, and Experiences."

Meta frames the release as an invitation rather than a finished product: it says it lacks the domain expertise to apply Rebalancer to fields like healthcare, energy, or logistics itself, and is counting on outside systems and optimization researchers to extend it. The code is on GitHub, with a Python package on PyPI and documentation covering the modeling API and the Explorer debugging UI.

Share this story

Stefan Holloway

Stefan Holloway covers programming languages, open-source ecosystems, and the CI/CD and API tooling that ships software for techshooked. He writes reproducibility-first, stating the version tested, showing the configuration, and separating a genuine workflow improvement from release-note marketing.