Tainted\\Coders

Bevy Archetypes

Bevy version: 0.19Last updated:

Bevy is an archetypal ECS. Archetypes enable efficient storage and processing of entities by grouping them according to their component composition.

An Archetype describes a unique combination of components. A world has only one Archetype for each set of components.

Starting without archetypes

Imagine we have a naive implementation of an ECS that stores components as arrays and entities as the index:

Entity IDHealthPlayerPositionEnemy
1100.0X(3, 2)
272.0(3, 1)X
340.0X(4, 2)
420.0(3, 0)X

This layout is naive because not all entities have all components. Notice how entities 1 and 3 have the same component layout, while 2 and 4 share a different one. These gaps from the missing components create problems for vectorized operations.

Vectorized means code that operates on entire arrays or vectors of data at once, rather than processing individual elements sequentially.

Vectorized operations can be performed efficiently using hardware optimizations like SIMD (Single Instruction, Multiple Data). These instructions allow simultaneous execution of the same operation on multiple data elements, often resulting in improved performance.

However, in the given scenario, because the rows have null values, it becomes challenging to write vectorized code that requires both arrays of Player and Enemy.

Vectorized operations usually rely on the assumption that corresponding elements in different arrays have the same index. If the entity ids in A and B are not aligned with the array indices, it becomes difficult to perform operations that depend on matching elements between the two arrays.

Improving the layout of our components

What if instead we stored the components how they are likely to be used?

Entity IDHealthPlayerPosition
1100.0X(3, 2)
340.0X(4, 2)
Entity IDHealthEnemyPosition
272.0X(3, 1)
420.0X(3, 0)

Now we have two tables, but each table forms a contiguous array of elements that can be optimized upon retrieval.

The two tables are each a different Archetype. One for the Player and one for the Enemy.

So we are storing our components in arrays of the same type, but inside of different tables representing a type (as in the sum of all its components) of an entity. We shall call this type: an Archetype.

What is inside an archetype?

An Archetype is a node that stores some metadata about the entities and components it contains, along with the edges that connect it to other archetypes:

// https://docs.rs/bevy/latest/bevy/ecs/archetype/struct.Archetype.html
pub struct Archetype {
  id: ArchetypeId,
  table_id: TableId,
  edges: Edges,
  entities: Vec<ArchetypeEntity>,
  components: ImmutableSparseSet<ComponentId, ArchetypeComponentInfo>,
  pub(crate) flags: ArchetypeFlags,
}

Archetypes are stored in a specific World and are locally unique.

// https://docs.rs/bevy/latest/bevy/ecs/archetype/struct.Archetypes.html
pub struct Archetypes {
  pub(crate) archetypes: Vec<Archetype>,
  /// find the archetype id by the archetype's components
  by_components: HashMap<ArchetypeComponents, ArchetypeId>,
  /// find all the archetypes that contain a component
  pub(crate) by_component: ComponentIndex,
}

Archetypes and tables share the same memory management. They both can only be created and are never destroyed. They persist until the World that owns them is dropped. Even the empty archetype (ArchetypeId::EMPTY) exists from the moment the World is created.

Creating archetypes

Spawning even a single entity with some set of components will create or update an archetype.

For example, if we spawn an entity with the following components:

world.spawn((
  Player,
  Health(100.0),
  Position { x: 0.0, y: 0.0 },
));

Bevy would create an Archetype for this entity with the components (one per column) inside a Table that looks like this:

PlayerHealthPosition
X100.0{ x:0.0, y:0.0 }

Each component is stored as a columnar array, with each row being one entity. The archetype's entities: Vec<ArchetypeEntity> maps each entity to its row in the table.

When a Query is used as a system parameter, it matches every archetype whose component list satisfies the query. Separately, Bevy statically inspects the component access each system declares to decide which systems can run in parallel.

When you add or remove a component on an entity you are moving that entity to a new archetype.

Adding a component the entity already has does not change its archetype, and a single insert can add several components at once when required components are involved.

Archetypes help find disjoint queries

A Component can be present in many archetypes. Bevy tracks this many-to-many relationship with Archetypes::by_component, a ComponentIndex that maps each ComponentId to the archetypes that contain it, along with an ArchetypeRecord describing how the component is stored there.

This relationship is also how the parallel scheduler figures out which queries can run at the same time.

At first glance you would think this shouldn't be able to run in parallel, since we need mutable access to the same Position component:

fn move_players(
  player_positions: Query<&mut Position, With<Player>>,
) {
  // ...
}

fn move_enemies(
  enemy_positions: Query<&mut Position, Without<Player>>,
) {
  // ...
}

If components were simply stored and accessed on one giant Vec<T>, then player_positions would have already borrowed the Position components and enemy_positions should not run in parallel. However, Bevy can prove statically that the two queries never touch the same entities.

Every system declares the component access of its queries as a FilteredAccessSet. When deciding whether two systems can run in parallel, the scheduler performs two checks:

  1. A coarse check of the raw component accesses. Two &mut Position queries fail this check.
  2. A fine-grained check of the With/Without filter sets. Because no archetype can both contain and not contain Player, the filters make the queries mutually exclusive, so the systems are considered disjoint and can run in parallel.

This works for both read-only and mutable queries, and it is a property of the access a query declares, not of which archetypes happen to exist at runtime.

When and how do archetypes get updated?

An Archetype comes into existence when we spawn or insert a Bundle (or when required components implicitly add other components). Given a set of components stored in either a Table or SparseSet or some combination of the two, the table is allocated first via Tables::get_id_or_insert. That TableId is then passed to Archetypes::get_id_or_insert, which either finds the matching archetype or creates a new one.

When we add or remove bundles from an Entity, it moves to a new archetype.

These moves can be quite expensive so Bevy caches these moves in Edges. Next time a bundle gets added we can fetch for a matching edge and find the target Archetype to move it to without computing it from scratch.

How does Bevy know our archetypes have changed?

Archetypes are never removed, so Archetypes is append-only. Because of this, Archetypes::generation can cheaply report a handle to the current highest archetype. It is simply ArchetypeId::new(self.archetypes.len()).

A query state stores the ArchetypeGeneration it has already processed. When the query is updated, Bevy iterates only over &archetypes[old_generation..] and checks whether each newly added archetype matches the query, so a query never has to re-scan the entire world.

How are archetypes different than tables?

Typically, each archetype has its own dedicated table. Archetypes only share tables when components are stored in sparse sets.

Table storage is columnar and optimized for iteration, while SparseSet storage is a HashMap-like mapping optimized for random lookup and regular insertion and removal.

Marking a component as SparseSet lets an entity gain or lose that component without moving its table components to a new table. If there was no SparseSet then archetypes and tables would be 1:1. So only in the case of a component being stored in a SparseSet are they ever potentially shared.

An archetype [A, B, C] and an archetype [A, B] will have the same table if C is a sparse component (stored in a SparseSet).