3.1 Anatomy of the eye and image formation⧉
The eye is an optical instrument, the biological counterpart of the camera assembled in Fundamentals of imaging: the same rays, focused by refracting surfaces onto a sheet of receptors. This chapter walks that hardware, from the cornea to the cortex, before the next chapters ask what the visual system computes from the image it forms.
Light entering the eye passes through the cornea and the lens, the eye's two refracting elements, the same Snell's-law optics as a camera lens (the cornea does most of the bending; the lens fine-tunes focus by changing shape, accommodation). The pupil, the iris's adjustable opening, is the eye's aperture. The image lands on the retina, a thin sheet of neural tissue lining the back of the eyeball (Figure 3.1.1).

The retina's light-sensitive cells come in two families. Rods are exquisitely sensitive and handle dim, night (scotopic) vision, but there is only one kind, so rod vision is colorless. Cones are less sensitive, work in daylight (photopic vision), and come in three types, the basis of color. Their distribution across the retina is wildly non-uniform (Figure 3.1.2). Cones are packed densely in the fovea, the tiny central pit we point at whatever we are scrutinizing; rods dominate the periphery. This is why you cannot read text out of the corner of your eye, and why faint stars are easier to see slightly off-center.
Physically these are minute cells. In the fovea the cones are the finest of all, only about $2\,\mu\text{m}$ across, packed at a center-to-center spacing near $2.5\,\mu\text{m}$ that sets the eye's acuity limit; toward the periphery cones fatten to $5$–$8\,\mu\text{m}$. Rods are slimmer still, roughly $1$–$2\,\mu\text{m}$ in diameter. A few micrometers is therefore the biological yardstick against which a camera's photosite is measured, and modern phone pixels have shrunk to below the size of a foveal cone (→ photosite sizes).
Two quirks of this layout matter. First, there are essentially no short-wavelength (S, "blue") cones in the very center of the fovea, and far fewer S cones than long- and medium-wavelength (L and M) cones overall. As a result blue is sensed at low resolution and is even slightly out of focus (the eye's chromatic aberration puts blue at a different focal plane), yet we perceive a crisp blue everywhere, because the brain fills it in. Second, luminance (our sense of brightness) is roughly the sum of the L and M cone responses, so green carries most of our spatial acuity. As we will discuss in Sensing color: multiplexing strategies, this is why camera color-filter mosaics (the Bayer pattern) use twice as many green cells as red or blue.
The retina is not passive and does substantial processing before anything leaves the eye. The photoreceptors do not report straight to the brain; they feed layers of intermediate cells that in turn drive the retina's output neurons, the ganglion cells, whose long fibers bundle together into the optic nerve, the cable to the brain. Already here the signal is reorganized: each output cell no longer reports raw brightness but compares a small spot to the ring of retina around it, a center-surround response that reacts to local contrast and edges rather than absolute level (the seed of the edge enhancement we return to under Spatial vision). From the eye the pathway passes through a relay station deep in the brain (the lateral geniculate nucleus) and on to the primary visual cortex at the back of the head, then to higher areas that handle motion, form, and recognition. The lesson for us: much of what we naïvely attribute to "the image" is really a construction, assembled in stages from a heavily pre-processed retinal signal.