Neural Cellular Automata, Explained
After reading this you will understand how a grid of identical cells, each running the same tiny network, can grow a texture from one seed and heal a hole you punch in it, and you will be able to reproduce the update by hand for a single cell.
What a neural cellular automaton is
A neural cellular automaton (NCA) is a grid of cells. Every cell holds a small vector of numbers: a red, green and blue channel, one "alive" channel, and a handful of hidden channels with no fixed meaning. Every cell runs the same update function on every step. There is no central controller. Whatever pattern you see is the fixed point (or the endless churn) of that one local rule applied everywhere at once.
The hook is repair. Seed a single living cell in the middle of the grid and the alive channel spreads outward until the pattern fills its region. Erase a disk of cells with a click and the frontier grows back inward. Nothing tracks the "target shape" from a global view. The healing is a side effect of a rule that says: a dead cell next to living neighbours comes alive, and a cell with no living neighbours dies.
This sandbox is honest about its limits. The alive channel follows a hand-designed logistic growth rule so every seed reliably grows and heals. The color and hidden channels evolve under a randomly initialised network, not a trained one. You are exploring the space of NCA dynamics, not watching a trained organism regrow a specific picture.
When to reach for this, and when not
An NCA is worth studying when you want an intuition for how local rules produce global structure, and how robustness to timing noise falls out of a stochastic update. It is a clean example of a model with shared weights across space (a convolution) whose behavior over time is the thing you care about, not a single forward pass.
It is the wrong lens for a few things. If you want a model that maps noise to a specific target image, a trained NCA can do that, but training it needs backpropagation through many time steps, which is too slow for a browser page. For learning generative modelling more broadly, the Diffusion Model Toy and the GAN Trainer (2D) show two other routes from noise to data, both trained live. If your interest is sequential generation from a learned distribution over tokens, the Markov Chain Text Generator is the simpler starting point.
The update rule and where every symbol comes from
Write the state of the whole grid at step t as S_t, a stack of channel images. One step has four parts: perceive, transform, mask, and stochastic add.
The perception stage is a fixed depthwise convolution. Each channel is filtered by three fixed kernels: identity (the cell sees itself), Sobel-x and Sobel-y (the cell sees horizontal and vertical gradients across its 3x3 neighbourhood). If you have C channels, perception produces 3C numbers per cell.
Here s is the cell's state vector, * is convolution over the 3x3 neighbourhood, and the three kernels are the fixed identity and Sobel filters. This stage has no learnable weights.
The perception vector p feeds a small dense network: a linear layer W_1 with bias b_1, a ReLU, then a second linear layer W_2. The output \Delta s is a proposed change to the cell state. In the original trained model, W_1 and W_2 are learned. Here they are set from the weight seed you pick.
The update is stochastic. Each cell draws a mask bit m \in \{0, 1\} that is 1 with probability equal to the fire rate. Cells whose bit is 0 keep their previous state. After the add, the alive mask sets any cell to zero if it has no living neighbour above the alive threshold. The Sobel filters see raw pixel gradients, which is how each cell knows whether it sits on an edge of the pattern or deep inside it.
One cell, one step, by hand
Use the demo defaults: one alive channel plus color, fire rate 0.5, alive threshold 0.1. Take a dead cell at the growth frontier. Its alive value is a = 0. Two of its eight neighbours are alive with alive values 0.9 and 0.4; the other six are 0.
- Check the mask: the maximum neighbour alive value is
0.9, which exceeds the threshold0.1, so this cell is allowed to grow. - Logistic growth rule for the alive channel. Let the neighbour alive sum be n = 0.9 + 0.4 = 1.3. The rule pushes the alive value toward \sigma(k(n - \theta)) with growth gain k = 4 and centre \theta = 0.5. Compute 4 \times (1.3 - 0.5) = 3.2, and \sigma(3.2) = 1 / (1 + e^{-3.2}) = 0.9608.
- Apply the stochastic mask. With fire rate
0.5, this cell fires roughly half the time. Suppose it fires this step. The alive value moves from0toward0.9608. With a step size of0.5the new value is 0 + 0.5 \times (0.9608 - 0) = 0.4804. - The color and hidden channels update by the random dense network at the same time, adding their own \Delta s scaled by
0.5.
The cell is now alive at 0.4804. On the next step it becomes a living neighbour for the cells beyond it, and the frontier advances by one ring. That is the whole engine of growth.
Why erased regions heal from the edge in
Alive masking is the reason. A cell with no living neighbour above threshold is forced to zero every step, no matter what the network proposes. When you erase a disk, every cell inside it is dead and, at the first step, only the cells on the rim of the hole have a living neighbour (from outside the hole). So only the rim can grow. The next step, that new rim has living neighbours one ring further in, and growth advances. The hole closes as a shrinking front, not all at once.
The number of steps to close a hole is roughly its radius in cells. A hole of radius 12 cells closes in about 12 to 24 update steps, the range coming from the fire rate: at fire rate 0.5 each frontier cell fires about half the time, so the front advances at roughly half speed and takes closer to the upper end.
Reading the stability metrics
The sandbox reports live metrics so you can classify the dynamics of a weight seed without staring. Three matter most.
- Mean absolute change
- The average of |\Delta s| across all firing cells this step. Near zero means the pattern has settled. A steady positive value means it churns forever.
- Channel saturation
- The fraction of cells whose color channels are pinned at 0 or 1. Above about
0.8the texture has blown out and detail is gone. - Living fraction
- The share of cells above the alive threshold. This should climb from near zero to a stable plateau as the seed fills its region.
A rough reading: mean absolute change under 0.01 with saturation under 0.5 is a stable texture. Mean change steady around 0.1 is a churning seed. Saturation racing past 0.9 in a few steps is a blow-up, common when the random W_2 has large entries.
Common mistakes when reading the sandbox
The first mistake is expecting a picture. Because the network is random, not trained, the color channels will not converge to any recognisable image. Judge a seed by its dynamics: does it settle, churn, or saturate? The alive channel is the only part designed to produce a clean, repeatable shape.
The second mistake is treating fire rate as a speed knob only. It also changes stability. A pattern that looks stable at fire rate 0.5 can start oscillating at fire rate 1.0 because every cell now updates in lockstep and small waves reinforce instead of averaging out.
Do not read a saturated grid as "converged". Saturation and convergence look similar (both stop changing) but mean opposite things. Convergence is mean absolute change near zero with saturation below 0.5. Saturation is the color channels pinned at their limits with detail destroyed. Always check both metrics before you call a seed stable.
The third mistake is raising the channel count and expecting richer behavior for free. More hidden channels give the random network more directions to grow in, which makes blow-ups more likely, not less, because the largest random weight has more chances to dominate. With a trained network the extra channels help; with random weights they mostly add instability.
How the fixed perception filters shape everything
The three fixed kernels decide what a cell can possibly respond to. Identity passes the cell's own value. The two Sobel kernels measure gradients, so a cell effectively knows two things about each channel: its level and its slope. A cell in the flat interior of the pattern sees near-zero Sobel outputs; a cell on an edge sees large ones. This is why the interesting action happens at boundaries, and why the healing front is sharp: the Sobel response is largest exactly where the pattern meets dead space.
Frequently asked questions
Why does the color look random when the shape is clean?
The shape comes from the hand-designed alive channel, which is the same for every seed. The color comes from a randomly initialised network that was never trained toward any target, so its output is whatever the random weights produce. Randomising the weight seed changes only the color and hidden dynamics.
Is this the same model as "Growing Neural Cellular Automata"?
The architecture matches: fixed identity and Sobel perception, then a small dense layer added to the state, with stochastic updating and alive masking. The difference is training. The original learns the dense-layer weights by backpropagation through time to grow a specific image. Here those weights are random and the alive channel is scripted, because full training is too slow for a browser.
Why is the update stochastic instead of synchronous?
Random per-cell firing stands in for biological cells that have no shared clock. It also forces robustness: any pattern that survives must hold up when neighbours update out of order. At fire rate 0.5 a frontier cell fires about half the time, which slows growth to roughly half speed but removes the risk of synchronized oscillation.
What fire rate should I use?
Start at 0.5. Push toward 1.0 to see faster growth and expose seeds that only look stable because they update in lockstep. Drop toward 0.1 to watch the frontier advance in slow, smooth rings.
Can I make a seed that never settles?
Yes. Randomise weights until the mean absolute change holds steady around 0.1 instead of falling toward zero. Those seeds churn indefinitely. They are as valid an outcome as the stable ones; there is no target the model is failing to reach.