draft · v0.1.241
💬Comments welcome. To leave a note, select any text and click the note / highlight button that pops up — or open the panel with the tab at the top-right (‹). Notes are visible only inside our private review group.
Computational Photography, an AI-powered Slopendium — 03 Visual perception and color
expand to📖 Full book outline1 parts · 6 chapters · 23 sections · 55 figures embedded · 6 placeholders · double-click a figure to enlarge
Part 3 VISUAL PERCEPTION AND COLOR
fig-light-journey
fig-light-journey · FUNDAMENTALS part opener — the journey of light: source → scene (reflection) → lens → sensor / retina → visual system, each step labelled with its chapter
This part picks up where [[Fundamentals of imaging]] leaves off: after light has been focused by a lens onto the retina or a sensor. Where part 2 treated the image as a physical measurement, this part asks what that measurement *means* to an observer. Color is not a property of light but a projection made in the eye and brain, and that single fact grounds white balance, color management, JPEG's chroma subsampling, and tone mapping downstream.
• **Anatomy of the eye and image formation** — the eye as an optical instrument: cornea, pupil, lens, and retina; the cone/rod mosaic and the fovea; and the eye's own refractive errors and how they are measured (inverse imaging).
• **Perceptual color and trichromatic vision** — why three cones make us trichromatic; color as the projection of a spectrum onto three numbers; metamers, color blindness, opponent processing, and color as linear algebra.
• **Measuring and encoding color** — the engineering payoff: measuring color (CIE), encoding it (linear / gamma / log, color spaces, CIELAB), and reproducing it (additive vs subtractive synthesis, gamuts, white balance, color management).
• **Sensing color: multiplexing strategies** — how a monochrome sensor is made to see color: the Bayer mosaic and demosaicking, 3-CCD prisms, stacked and dispersive (nano-prism / metasurface) sensors, and the discrimination-versus-hardware trade each makes.
• **Visual processing** — what the visual system does with light once it has it: light adaptation, lightness and color constancy, contrast, spatial and temporal vision (the contrast sensitivity functions), attention, and the constructed nature of perception.
• **Animal eyes** — human vision is one solution among many: how other eyes form images and sense color, from compound eyes to tetrachromats to the mantis shrimp.
3.1 Anatomy of the eye and image formation
fig-eye-cross-section
fig-eye-cross-section · eye cross-section 🟨
fig-retina-layers
fig-retina-layers · retina layers & photoreceptor synapses 🟨
fig-cone-rod-distribution
fig-cone-rod-distribution · cone/rod distribution & fovea 🟨
fig-visual-pathway
fig-visual-pathway · pathway retina→LGN→V1 🟨
fig-photoreceptor-cell
fig-photoreceptor-cell · anatomy of a rod & cone cell 🟨
fig-rods-cones-micrograph
fig-rods-cones-micrograph · micrograph of rods & cones (primate retina) 🟨
• light, lens, pupil, retina,
• cones/rods, distribution, fovea — note there are **no S (blue) cones in the very centre of the fovea**, and far more L+M than S overall, so **blue is low-resolution and a bit out of focus** (chromatic aberration) — yet we perceive blue everywhere (the brain fills it in). Luminance ≈ the **L+M cone sum**, which is why **green carries most of the acuity** (and why Bayer uses 2× green).
• cells, ganglions, on-retina spatial processing, visual cortex, V1, higher.
• **the eye as a camera — and its refractive errors.** Like any lens system the eye can mis-focus: **myopia** (eyeball too long / cornea too curved → image forms *in front* of the retina), **hyperopia** (too short → *behind*), **astigmatism** (a non-spherical cornea → two focal lines, not one point), and **presbyopia** (the aging lens stiffens and loses **accommodation**). Each is corrected by a lens quoted in **diopters** (D = 1/focal-length in metres): a negative/diverging lens for myopia, positive/converging for hyperopia, a **cylindrical** lens (with an axis) for astigmatism. [`fig-refractive-errors`]
• **how the prescription is *measured* (the machines), and why it's just inverse imaging.** **(1) Retinoscopy** — shine a moving streak of light into the eye and watch the **retinal reflex** sweep across the pupil; its direction and speed reveal myopia vs hyperopia, and trial lenses are stacked until the reflex stops moving ("neutralization"). **(2) Autorefractor** — automates this objectively: it projects an **infrared** target onto the retina and analyses the **returning image** (its focus/size, or a Hartmann–Shack / Scheimpflug pattern) to solve for the corrective lens that makes the retinal image sharp — essentially a little camera running autofocus *on your eye*, in seconds. **(3) Phoropter** — the familiar "better one… or two?" wheel of trial lenses over an eye chart: a **subjective** refinement of the machine's objective estimate. **(4) Wavefront aberrometer (Shack–Hartmann)** — for the full optical picture: a **lenslet array** samples the wavefront leaving the eye and measures **higher-order aberrations** (coma, spherical, trefoil), not just sphere/cylinder — the same wavefront sensing used in **adaptive optics** and in planning **LASIK**. The throughline: measuring an eye is **inverse imaging** — find the optics that would make the retinal image (or the returning wavefront) ideal. (Cross-ref wavefront / Shack–Hartmann sensing and adaptive optics in [[Optics, lenses, and aberration correction]].)
3.2 Perceptual color and trichromatic vision
3.3 Measuring and encoding color
3.4 Sensing color: multiplexing strategies
fig-color-multiplexing
fig-color-multiplexing · four colour-sensing multiplexing strategies: temporal, spatial (Bayer), **dichroic 3-CCD** beam-split, depth (Foveon, with realistic broad/overlapping layer colours + colour-matrix note) 🟨
⬜ figure not yet created
temporal multiplexing artifact — Prokudin-Gorskii color fringes on moving water fig-prokudin-gorskii
• four ways to turn a monochrome sensor into a color one — **multiplex** the three (or more) measurements across some axis:
• **in time**: shoot R, G, B frames sequentially through filters — Maxwell's first color photograph (1861), **Prokudin-Gorskii**, flatbed scanners, astronomy filter wheels. Cheap and full-resolution, but **fails on motion** (the classic color fringes on moving water).
• **in space (the dominant choice)**: a **color filter array (CFA)** — the **Bayer mosaic** (2 green : 1 red : 1 blue) — one color per photosite, then **demosaick** to fill the rest; the eye does the same with its interleaved cone mosaic. Variants: Fuji's **X-Trans** (a larger, less-periodic tile to fight moiré) and **complementary CMY/CYGM** CFAs (more light through, messier color). Needs an **optical anti-aliasing (low-pass) filter** to tame high-frequency color artifacts, and suffers some channel **crosstalk**.
• **beam-splitter (prism)**: a **dichroic prism** splits the rays onto **three sensors** — **3-CCD / 3-chip**, the classic **broadcast video camera** — wasting no photons at full per-channel resolution, but bulky, costly, and hard to align.
• **in depth (stacked)**: stack wavelength-selective layers so one location senses all three — **Foveon** (silicon absorbs longer wavelengths deeper; used in **Sigma** cameras), echoing **Kodachrome**'s stacked dye layers. Full color at every pixel (**no demosaicking**), but trickier color separation and more chroma noise.
• **spectral / Lippmann**: record the standing-wave **interference** of the full spectrum (Lippmann, 1908 Nobel; a precursor to holography) — true spectral capture, not just 3 numbers.
• **hybrid**: combine these (e.g. spatial + temporal in video).
• this is *analysis* (sensing color) — the mirror of color *synthesis* (reproduction, above); the **demosaicking** algorithm itself is deferred to Basic Image Processing (which is why this is just the *sensing* half)
• **trichromatic (perceptual) vs spectral (physical) capture**: almost all imaging systems aim to capture **perceptual, trichromatic color** — three numbers that reproduce what a *human* would see (RGB ≈ the cones), so metamers are, by design, indistinguishable. **Multispectral / hyperspectral** imaging instead samples the **physical spectrum** in many narrow bands (tens to hundreds), keeping distinctions the eye throws away — for remote sensing, agriculture, art conservation, and machine vision, where the *material*, not the *appearance*, is what matters (and metamers must be told apart). Lippmann (above) is the limiting case: the full continuous spectrum.
3.5 Visual processing
• the perceptual machinery the visual system runs *on top of* color: adapting to level, discounting the illuminant to report surface reflectance (constancy), working in ratios (contrast), and resolving detail unevenly across spatial/temporal frequency (the CSFs) — the properties compression, tone mapping, and display design optimize for.
3.6 Animal eyes
fig-eye-evolution
fig-eye-evolution · the evolution of the eye, in cross-sections: flat photoreceptor patch (no image) → cup (directional) → pinhole (an image, no lens) → lensed eye — recapitulating sensor → camera-obscura → lens (Bonus: animal eyes) 🟨
• **the eye evolved as a sequence of small improvements**, and it recapitulates the optics of this book — from "no image" to a focused one. Following Nilsson's stages (Land & Nilsson, *Animal Eyes*):
• **a flat patch of photoreceptors** (an *eyespot*): just **photosites**, **no image formation** — it tells light from dark and, crudely, which side has *more* light (phototaxis), but cannot form a picture. This is the bare **sensor** with no optics.
• **a concavity / cup**: fold the patch into a **pit** and the rim's shading gives **directional** sensitivity — the deeper the cup, the better you can tell *where* light comes from.
• **a pinhole**: close the cup almost shut → a **pinhole eye**, a real (if dim) **image** with no lens at all (the **nautilus** still uses one) — exactly the camera obscura of [[#Pinhole image formation and linear perspective]].
• **a lens**: fill the aperture with a **refractive lens** (often a graded-index sphere) → a **bright, focused** image, independently evolved in fish, cephalopods, and vertebrates. Add a **cornea, iris, and accommodation** and you have the human eye.
• so the human eye is *one* solution; evolution reached the same goal by very different routes (below).
• **a tour of those other solutions** — a useful contrast to the human eye and to camera design:
• **compound eyes** (insects): many **ommatidia** → wide field of view and fast temporal response, but low spatial resolution
• **more cone types**: birds / reptiles are **tetrachromatic** (a 4th cone, into the UV); the **mantis shrimp** has ~12–16 photoreceptor classes (yet poor color *discrimination* — it seems to recognize rather than compare) — the **divergence of opsins** across animal groups is the color sidebar in [[#Color]]
• **tapetum lucidum** (cats, deer): a reflective layer behind the retina that boosts night vision (and causes eye-shine)
• **polarization vision** (cephalopods, many insects): they see light's **polarization** — a channel we're blind to
• **other eye designs**: **pinhole** (nautilus), **mirror eyes** (scallops — image formed by a concave mirror, not a lens), and the front-facing (**stereo**) vs side-facing (**field-of-view**) placement trade-off
• further reading: Land & Nilsson, *Animal Eyes*; the "Evolution of Eyes" chapter in *Vision* (Cambridge); a friendly overview at phos.co.uk, "The Evolution of Sight".
3.6.1 The optics: many ways to form an image
3.6.2 The color vision: opsins remixed