Why this note exists
The Chapter 9 generator’s vocabulary is not invented; it is measured off the census. This note records that vocabulary and the descriptive contract under which the generator uses it.
1. The kit and its conditional rules
Ten data-derived geometric primitives (size band × aspect band × rectangular/complex) cover 90.7 per cent of all 12,849 spaces (top three 65.2 per cent). The single largest class is the undimensioned-rectangle placeholder (41.7 per cent): not a true geometric primitive but the spaces without printed dimensions; restricting to the 7,245 spaces with a resolved dimensioned shape, the ten most common dimensioned primitives cover 88.7 per cent of them. Either way the generative kit is small: a whole dwelling is drawn from about seven distinct primitives (kit-reuse ≈ 2.4). Each primitive carries a size signature (for example M-square-rect ≈ 3.6 × 3.2 m, bedroom-dominated; L-oblong-rect ≈ 5.5 × 3.7 m, living; XL-square-rect ≈ 6.0 × 5.8 m, garage). The generator selects and sizes a primitive for a context using the conditional design rules of Appendix H.5 (the empirical distribution of size for a given category × host × graph-role × bedroom-count cell).
2. The generator is descriptive: a lookup-and-redraw sampler, not a fitted model
The conditional rules are empirical conditional distributions tabulated over the complete census. The generator’s primary output for a context is the observed band (median and IQR) a designer dials within: a deterministic lookup. Where a single value is wanted, it is redrawn from the literal observed multiset of that cell, so the generator can emit nothing the corpus did not exhibit in that context; it never synthesises an intermediate value through a fitted curve, and it makes no inferential claim, because over a complete enumeration the band is the population’s own value, not an estimate.1 This is categorically unlike procedures that manufacture pseudo-samples to estimate sampling variability (none is used) and unlike learned generators that fit weights to a training sample.
3. Two hard constraints on the generator
- Out-of-corpus guard (required, mechanised). For any context cell the census never observed, the generator flags and refuses to synthesise rather than interpolating a size; interpolation across cells would be inferential. This must be a hard check in the generator contract, not a discipline.
- Coupling stays on the topology graph. The adjacency vocabulary uses the opening grammar (design rule DR-4) for opening type, but the obligatory-versus-preferred coupling split (49 HARD / 50 SOFT) is read from the canonical topology graph (the 15,957-zone access basis), because the geometry graph realises only about 272 of roughly 680 category pairs and cannot carry the coupling contract.
Thin context cells are reported with their exact observed count and spread, with the wider parent-type band carrying the recommendation: a transparent dial-stability convention, not a sample-size correction.
4. Positioning
The kit extends the shape-grammar tradition by offering a measured generative vocabulary that is sampled rather than an author-specified rule set that is enumerated;2 and it differs from learned floor-plan generators (for example the RPLAN/Graph2Plan and House-GAN families) by exposing explicit, inspectable conditional distributions instead of design knowledge locked inside network weights.3 Descriptive throughout; CANDIDATE until operator MR-13.
Notes
- The methodological point is standard for a complete enumeration: S. Gorard, Research Design (London: SAGE, 2013), p. 54. ↩︎
- G. Stiny and J. Gips, “Shape Grammars and the Generative Specification of Painting and Sculpture,” Information Processing 71 (IFIP, 1972): 1460-1465. ↩︎
- R. Hu et al., “Graph2Plan: Learning Floorplan Generation from Layout Graphs,” ACM Transactions on Graphics 39(4) (2020); N. Nauata et al., “House-GAN,” ECCV 2020. ↩︎