The Exoplanet Catalogue2 of 3

Plausible Worlds from Sparse Parameters

Procedural appearance synthesis for planets no one has seen, and where measurement ends and fiction begins

Method description and epistemic boundsJuly 202628 min read

First published at exoplanet-catalogue.vercel.app.

Abstract

No exoplanet has ever been resolved as more than a point of light. What is known about any of them is a short vector of numbers: orbital period, semi-major axis, eccentricity, radius and/or minimum mass, and the host star's temperature, radius and mass. Yet the public image of the field is one of worlds, oceans, cloudscapes, lava plains. Someone paints those pictures, and the relationship between the paint and the numbers is usually undisclosed.

This paper describes, and then deliberately bounds, a system that makes that relationship explicit. A classification pipeline maps each catalogue record to one of sixteen physically motivated planet types through a documented decision tree grounded in the characterisation literature: the Fulton radius gap separates rocky worlds from sub-Neptunes, bulk density separates water worlds from rock, the Kopparapu habitable-zone flux polynomials position temperate worlds between Mars-like and Venus-like end states, Sudarsky's classes colour the giants, and a spectral-type-dependent heuristic flags tidal locking, yielding the distinctive "eyeball" regimes of locked ice-ocean and lava worlds. A procedural shader layer (~2,100 lines of GLSL, ~69 uniforms) then renders each type from seeded noise: Perlin fractional Brownian motion, ridged and cellular variants, and domain warping, under a wrap-diffuse lighting model with a fresnel-mixed atmosphere approximation.

The paper's second half grades every visual decision on a four-tier speculation gradient, measured, derived, theory-guided, fictive, and states the ceiling plainly: the output is plausibility, not prediction. A rendered continent is fiction constrained by physics; its shoreline is a hash function. We argue that this parameterised, deterministic, auditable fiction is an honest middle path between data-blind illustration and data-only tables, and we specify the audit that would test it.

1. Introduction

1.1 The epistemic situation

It is worth being precise about how little seeing has occurred in exoplanet science, because the whole design follows from it. The overwhelming majority of confirmed planets were detected by transit photometry (a periodic fractional dimming of the host star, yielding a radius ratio) or radial velocity (a periodic Doppler shift, yielding a minimum mass); in both cases the planet itself contributes no image at all. Direct imaging exists, and HR 8799's four giants are its canonical family portrait (Marois et al. 2008), but each such image is an unresolved point spread function: a location and a brightness, not a disc, not a surface.

Everything beyond the point is inference, and the inferences are genuinely impressive. Thermal phase curves have yielded longitudinal brightness maps of hot Jupiters (Knutson et al. 2007; Majeau, Agol & Cowan 2012) and even of a lava super-Earth (Demory et al. 2016); occultation photometry produced a measured colour for one planet, HD 189733 b, which is deep blue, most likely from silicate scattering rather than oceans (Evans et al. 2013); transmission spectroscopy has detected water, sodium, carbon dioxide, and the muting signature of clouds and hazes (Kreidberg et al. 2014; Sing et al. 2016; JWST Transiting Exoplanet ERS Team 2023; Madhusudhan 2019). These are real constraints on appearance. They are also, for a visualiser, desperately sparse: a colour for one planet, a one-dimensional temperature map for a few dozen, molecular inventories for perhaps a hundred, and for the thousands that remain, nothing but the orbital and bulk parameters.

The field's public imagery has responded with the artist's impression: skilled, often scientifically advised, and, as a genre, unlabelled at the pixel level. A viewer cannot tell which features of a given rendering are constrained and which are composition. This paper describes the alternative taken by the exoplanet-catalogue project (the system is described in the companion paper): replace the artist with an explicit function from catalogue record to rendered world, and then document exactly what that function does and does not know.

1.2 Contributions

  1. A documented classification function from sparse, partially missing catalogue records to sixteen renderable planet types, with every threshold stated and sourced to the characterisation literature, and every missing-data fallback explicit.
  2. A parameterised appearance layer in which each type is rendered by seeded procedural shading, no textures, no per-planet art, such that appearance is a pure, reproducible function of the record.
  3. The speculation gradient: a four-tier grading (measured, derived, theory-guided, fictive) applied to every class of visual decision the system makes, offered as a reusable disclosure framework for scientific visualisation beyond this project.
  4. A statement of the ceiling, plausibility rather than prediction, together with the audit design that would test whether the plausibility claim itself holds.

2. The classification layer

2.1 Inputs and fallbacks

The classifier receives, per planet: mass and radius (Jupiter units, either or both possibly absent), semi-major axis, eccentricity, and the host star's temperature, mass and radius. Sparsity is handled by a documented fallback chain rather than by rejection.

Two derived quantities drive the tree. Effective stellar flux, S_eff = (T★/5780 K)⁴ · (R★/R☉)² / (a/AU)², is Stefan–Boltzmann luminosity normalised to Earth's insolation (calibration: Venus ≈ 1.9, Mars ≈ 0.43). Equilibrium temperature, T_eq = T★ √(R★/2a), is the standard zero-albedo estimate. And one derived flag: tidal locking is assumed when the orbit is close (inside 0.15 AU for M dwarfs, 0.05 for K, 0.03 for G, 0.02 for hotter stars), near-circular (e < 0.25), and the planet is rocky-sized, a coarse proxy for the spin-synchronisation timescales treated properly by Barnes (2017).

2.2 The decision tree

The tree branches first on size, then on energy and composition. Verbatim in its thresholds:

The eyeball regimes deserve their citation trail, since they are the system's most exotic-looking and best-grounded output. A tidally locked temperate world plausibly freezes on its permanent night side and thaws only near the substellar point, the "eyeball" configuration of climate modelling (Joshi, Haberle & Reynolds 1997; Pierrehumbert 2011), popularised under exactly this name (Raymond 2015). The renderer's locked ice-ocean world (open ocean at the substellar point, pack ice and floes toward the terminator, a solid cap on the far side) and its locked lava world (molten dayside pool, graduated to a dark frozen crust) are direct visual transcriptions of that literature.

Figure 1. TRAPPIST-1 f rendered as a tidally locked ice-ocean eyeball: dark open ocean on the star-facing hemisphere, Voronoi-cracked pack ice consolidating toward the night side. Classification from the recorded record (rocky size, habitable-zone cold edge, locked); every surface feature is seeded noise.
Figure 1. TRAPPIST-1 f rendered as a tidally locked ice-ocean eyeball: dark open ocean on the star-facing hemisphere, Voronoi-cracked pack ice consolidating toward the night side. Classification from the recorded record (rocky size, habitable-zone cold edge, locked); every surface feature is seeded noise.

2.3 Within-type positioning

Classification is not the end of the data's influence. Temperate worlds are positioned within the habitable zone by the Kopparapu polynomials evaluated for the actual host star: the planet's S_eff is normalised between the computed runaway-greenhouse and maximum-greenhouse fluxes, nudged by a logarithmic mass term, and the result interpolates every surface parameter, sea level, ice-cap extent, cloud cover, atmospheric rim colour, along a three-anchor spline from a Mars-like state through an Earth-like state to a hot, humid end state. Giants are similarly modulated: Saturn-like pallor below half a Jupiter mass, deepened banding above 1.5, a washed-out "puffy" blend for inflated radii, following the qualitative trends of the irradiated-giant literature (Sing et al. 2016). Host-star temperature selects among curated surface palettes, so that red-dwarf worlds render under redder insolation than F-star worlds, and a per-planet seed applies bounded hue jitter so that sibling planets of one type remain distinguishable.

3. The appearance layer

3.1 Noise machinery

All surfaces derive from analytic noise; there are no image textures. The foundation is classic 3D Perlin gradient noise (Perlin 1985, 2002; GLSL implementation after McEwan et al. 2012), wrapped into fractional Brownian motion families: a high-detail fBm at eight octaves (three at low level-of-detail), a ridged variant (folded and fourth-powered, for mountain chains and flow channels), a sinusoidally modulated "cloud" variant, and a 3×3×3-cell Voronoi/cellular basis (Worley 1996) for craters, ice-floe plates and lava crust. Continents, gas-giant bands and ice boundaries are all domain-warped, noise evaluated at coordinates displaced by other noise (Quílez 2002–2008; the two-stage warp on the giants follows the canonical construction), which is what moves the output from static on a sphere toward geology. The star surface runs a separate 4D simplex fBm (five octaves) so that granulation animates in time; a temperature-derived tint applies an approximate blackbody colour continuously from M-dwarf orange to O-star blue.

3.2 The regime shaders

Seven fragment programs cover the sixteen types.

Figure 2. Eight featured planets as the pipeline classifies and renders them from their catalogue records: an ice-ocean eyeball, two temperate worlds, Sudarsky Class IV and V hot Jupiters, a frozen pulsar planet, a Class II water-cloud giant, and a lava eyeball.
Figure 2. Eight featured planets as the pipeline classifies and renders them from their catalogue records: an ice-ocean eyeball, two temperate worlds, Sudarsky Class IV and V hot Jupiters, a frozen pulsar planet, a Class II water-cloud giant, and a lava eyeball.

Gas giants

The giant program advects its noise field with a latitude-dependent Coriolis rotation (fast equator, slow poles), lays banding as a sine of warped latitude whose frequency is a per-type uniform, seeds a handful of anticyclonic storm vortices with compressed, cubic-falloff eyes, roughens band boundaries with edge turbulence, and darkens poles and storm cores. Class V giants add thermal emissive glow; ice giants swap to methane blues with faint banding, a Neptune-like dark-spot pair, and their own limb darkening.

Terrestrial worlds

The terrestrial program builds a continent function from three domain-warped fBm layers with an S-curve contrast push around a sea level set by classification; ocean colour grades through deep, mid and shallow water with a coastal shelf; land is zoned by latitude and a moisture field into wet and arid lowlands, highlands, tundra, exposed rock and snow-capped peaks; four scales of procedural bump normals (terrain gradient, ridged chains, hills, micro-detail) light the relief; polar caps grow with domain-warped, noise-jittered edges. Ocean pixels receive a specular lobe (Blinn-style, exponent 32) gated by the terminator.

Locked eyeball worlds

The locked regimes replace latitude zoning with substellar-angle zoning: the ice-ocean case keeps open water under the star, then slush, then pack ice with Voronoi crack networks (dark seawater in the cracks), then a solid cap on the night side; the lava case renders a molten pool graded through a seven-stop temperature ramp toward a dark basalt night side, with lava rivers along major Voronoi edges and residual-heat seams on the cold hemisphere.

Cloud layers

Clouds are a separate translucent sphere (radius 1.006, temperate and water worlds only) with four superposed systems, domain-warped cumulus masses, stretched cirrus wisps, ridged frontal boundaries and fine convective texture, together with an ITCZ band at the equator and Hadley-like banding at a per-type frequency; on locked worlds the entire layer reorganises into a logarithmic-spiral hurricane centred on the substellar point, with a cleared eye, feathered arms, and a convergence band at the terminator. Coverage is thresholded, so classification tunes cloudiness; edges are eroded by noise; and the layer fades out across the terminator.

Hazy, airless and frozen worlds

The remaining programs cover hazy worlds (sub-Neptunes and Venus analogues: thick domain-warped cloud tops, slow rotation) and airless or frozen rocky worlds (ridged terrain with four scales of Voronoi cratering; lava variants swap craters for flow-warped noise, emissive in the lows, and on locked hot rocks a day–night glow gradient from 8 % to 200 %).

3.3 Lighting and atmosphere

Lighting is deliberately a stylised model, not radiative transfer. The diffuse term is wrap lighting, light = clamp(((N·L)·w + w)^p, 0, 1), a long-standing real-time softening of the terminator (the half-Lambert family; Mitchell, McTaggart & Green 2006), with wrap w and power p exposed as uniforms and tuned project-wide (w = 0.65, p = 4). A fresnel-squared rim term stands in for limb ambience, scaled by an ambient uniform whose default is zero: night sides are dark. The atmosphere is a single fresnel-mixed shell whose rim colour interpolates from a twilight tone to a day tone across the terminator, masked by the same wrap function; this is an intentionally cheap approximation in place of physically based scattering (Nishita et al. 1993; Bruneton & Neyret 2008), and it is one of the paper's clearly flagged fidelity ceilings.

Determinism is total. Given one catalogue record and one seed, every octave of every noise call is reproducible; the same world renders on every machine, in the live viewer and in the baked thumbnails alike (companion paper, §3.7).

4. The speculation gradient

The system's claim rests on being able to say, for any visible feature, which kind of statement it is. We grade four tiers.

Tier 1: measured

Orbital period, semi-major axis, eccentricity, radius and/or minimum mass, stellar temperature, radius and mass, distance. These come from the catalogue with published uncertainties, which the renderer does not yet surface, a stated limitation. Visual consequences: orbital motion, relative sizes, star colour.

Tier 2: derived

Bulk density, S_eff, T_eq, habitable-zone position. Deterministic arithmetic on tier 1, inheriting its gaps: T_eq assumes zero albedo, and density inherits the mass–radius fallback where one input is missing. Visual consequences: which side of each classification threshold a planet falls.

Tier 3: theory-guided inference

The planet types themselves. That a 1.4 R⊕ planet at S_eff 0.9 is "temperate", that a ρ < 2 world is volatile-rich, that a close-in rocky planet is locked, that a 1600 K giant carries silicate clouds: each is a defensible reading of the literature (Kopparapu 2013; Fulton 2017; Zeng 2019; Barnes 2017; Sudarsky 2000, 2003), and each could be wrong for any individual planet. The observed diversity of the population, clear and cloudy hot Jupiters at the same temperature (Sing et al. 2016), guarantees a nonzero per-planet error rate that no threshold tree can remove. Visual consequences: everything categorical, ocean or lava, banded or blue.

Tier 4: fictive

Continent shapes, shoreline fractality, cloud configurations, storm placements, crack networks, palette jitter. These are hash functions. They are constrained fiction, a temperate world's fiction drawn from Earth-like morphology, a locked world's from eyeball climate states, but no pixel of them is knowledge. The system's one unbreakable rule is that tier 4 must never be adjusted per planet by hand: the moment a specific world's continents are art-directed, the pipeline's claim to be a function of the data is forfeit.

Stated as a ceiling: the renderer offers plausibility, not prediction. It is a visual hypothesis generator whose hypotheses are typed, sourced and reproducible, closer to a climate-model schematic than to a photograph, and closer to a photograph than to a table. The appropriate reading of any rendered world is a planet of this measured kind could look like this, and never this planet looks like this.

This grading also clarifies the relationship to the artist's impression. The difference is not talent (a skilled artist encodes more atmospheric science in a single matte painting than this shader knows) but auditability: here the mapping from record to image is code, versioned, uniform across four thousand systems, and gradable tier by tier. The genre-level argument is continued in the third companion paper.

5. Related work

Planetary procedural generation

The techniques are graphics-canonical: gradient noise and its fBm assemblies (Perlin 1985, 2002; Musgrave et al. 1989; Ebert et al. 2003), cellular bases (Worley 1996), GLSL noise implementations (McEwan et al. 2012), domain warping (Quílez). Whole-planet procedural systems are mature in entertainment; Space Engine renders a procedurally infinite universe seeded by real catalogues where available, and No Man's Sky (Hello Games 2016) generates worlds at galactic count. But their generators optimise variety and wonder, not fidelity to a specific record, and neither documents a per-feature epistemic grading. Scientific-visualisation systems (Eyes on Exoplanets; OpenSpace, Bock et al. 2020) conversely tend to use stock class imagery rather than per-record synthesis. The niche claimed here is narrow but real: per-record, literature-thresholded, tier-graded procedural appearance over a complete public catalogue.

Exoplanet characterisation

The classification leans on the mass–radius and demographics literature (Weiss & Marcy 2014; Rogers 2015; Fulton et al. 2017; Chen & Kipping 2017; Zeng et al. 2019), the habitable-zone formulation of Kopparapu et al. (2013) after Kasting et al. (1993), giant-atmosphere theory (Sudarsky et al. 2000, 2003; Madhusudhan 2019), locked-world climatology (Joshi et al. 1997; Pierrehumbert 2011), and the direct appearance constraints cited in §1.1. The renderer contributes nothing to that literature; it is a consumer with citations.

6. Evaluation design

Not yet run; specified for the record.

Classification audit

Sample ~100 catalogue planets stratified across the tree's branches, including all with published characterisation (TRAPPIST-1 b–h, GJ 1214 b, 55 Cnc e, HD 189733 b, the HR 8799 giants, known super-puffs); have two exoplanet researchers independently assign types from the same input vector; report agreement with the tree per branch (Cohen's κ), and, the more interesting number, the cases where the literature contradicts the tree (GJ 1214 b's flat spectrum, for example, demands cloud or haze cover that density alone would not assign; Kreidberg et al. 2014).

Perceptual honesty

Show readers rendered worlds with and without a tier-graded legend ("measured / inferred / illustrative"); measure whether the legend shifts stated confidence about specific features ("does this planet have continents?" should move from majority-yes toward majority-unknown). This is the direct test of the paper's central claim, that graded speculation can be communicated and not merely documented.

Determinism regression

A pixel-hash test across GPUs and drivers for a fixed record set, since the reproducibility claim ("same record, same world") is in practice bounded by floating-point and driver variance; the deviation should be measured, not assumed zero.

7. Limitations

The tree is thresholds, not posteriors

A planet at S_eff 1.05 renders Venus-like; at 1.03, temperate. Real inference would be probabilistic in the measurement uncertainties, which the catalogue provides and the classifier ignores. Rendering the modal world where the data straddle a boundary overstates confidence exactly there.

T_eq ignores albedo; the locking heuristic ignores time

Both are acknowledged coarse proxies (§2.1); both gate categorical visual outcomes.

Earth-centrism of the fictive tier

Cloud systems are Earth's (ITCZ, Hadley bands, cyclonic spirals); temperate palettes are Earth biomes hue-shifted by stellar class. The fiction is drawn from a sample of one inhabited atmosphere, and the true morphological diversity of temperate exoplanets is unknowable from it.

No radiative transfer, no spectra

Colours are palette assignments guided by the literature, not forward-modelled from atmospheres; the one planet with a measured colour (HD 189733 b) is matched by class, not computed.

Sparse records render with undiminished confidence

The interface does not yet distinguish a fully characterised planet from an m sin i-only detection, a limitation shared with the companion paper; the "what is this picture based on?" affordance is the highest-priority roadmap item this analysis motivates.

The plausibility claim is untested

Until the audit and perception studies run (§6), "plausible" is the author's judgement with citations, not a measured property.

8. Conclusion

A planet in a catalogue is a vector of numbers; a planet in the imagination is a place. Every exoplanet visualisation bridges that gap somehow, and the bridge is usually invisible. This paper has described a bridge built to be inspected: a classification function with stated thresholds and sources, an appearance function with stated machinery and seeds, and a grading that assigns every visible feature to measured, derived, theory-guided or fictive status. The construction is bounded in the sense its ceiling names: plausibility, not prediction. No shader will tell us what TRAPPIST-1 e looks like. What a shader can do, done this way, is show four thousand data points as the kinds of worlds the evidence permits, while keeping the receipt for every choice. The wider question, whether such bounded showing serves public understanding better than the alternatives, is taken up in the companion survey.

Acknowledgements

The classification stands entirely on the exoplanet characterisation literature cited throughout; the shading stands on four decades of procedural graphics beginning with Perlin. Errors of reading in either direction are the author's.

References

Exoplanet characterisation and demographics

Procedural graphics and real-time shading

Systems

Appendix A: The types at a glance

TypeTrigger (verbatim thresholds)Visual treatment
Hot Jupiter Vgiant, T_eq > 1400 Ksilicate tans, thermal emissive
Hot Jupiter IVgiant, T_eq > 900 Kalkali near-black
Warm giantgiant, T_eq > 350 Kcloudless azure
Cool giantgiant, T_eq > 150 Kwater-cloud white
Cold giantgiant, elseammonia bands, Jupiter-like
Ice giant3.5 < R ≤ 6 R⊕methane blue, dark spots
Lava eyeballlocked, S_eff > 25, ρ ≥ 3dayside melt pool, basalt night
Lava worldS_eff > 25, ρ ≥ 3global magma, emissive lows
Hot rockyS_eff > 1.04, dense/smallMercury greys, cratered
Venus-likerocky, S_eff > 1.04, R > 0.5 R⊕thick warped haze
Ice-ocean eyeballlocked, ρ < 2 or HZ cold edgesubstellar ocean, floe rings, night cap
Water world1.75–3.5 R⊕, ρ < 2deep ocean, sparse islands
Sub-Neptune1.75–3.5 R⊕, ρ unknown/midblue-grey haze
Temperaterocky, 0.35 < S_eff ≤ 1.04Mars↔Earth↔hot-humid lerp by HZ position
Frozenrocky, S_eff ≤ 0.35pale ice, ridged
Unknown(defined, unreachable)

Appendix B: Sourcing notes

  1. The 0.27 mass–radius exponent is used here as an engineering constant in the terran regime; Chen & Kipping's (2017) fitted terran exponent is ~0.28 with stated uncertainty, and the correspondence is approximate by design.
  2. The tidal-locking semi-major-axis cutoffs (0.15/0.05/0.03/0.02 AU by stellar class) are a legibility heuristic, not a fit to Barnes (2017); locking is a timescale, not a boundary.
  3. "Eyeball planet" is a popularising term (Raymond 2015) for configurations studied more soberly as synchronously rotating climates (Joshi et al. 1997; Pierrehumbert 2011); the renderer's iconography follows the popular term knowingly.
  4. Wrap-lighting values (w = 0.65, p = 4) and all palette hexes are aesthetic tuning, tier 4 in this paper's own grading, and are recorded in version-controlled defaults rather than claimed from any source.