Chapitre 1
The Visual Brain: How We See and Think Through Images
When you look at this page, what do you actually see? If you believe you're taking in every word, every detail, every element all at once, you're experiencing what scientists call a grand illusion. In a remarkable experiment, researchers asked pedestrians for directions while secretly switching one researcher with another mid-conversation. Astonishingly, over half the participants failed to notice they were suddenly speaking to a completely different person. This wasn't because their eyes weren't working-it was because their brains were focusing only on what mattered for the immediate task: giving directions.
Colin Ware's groundbreaking work "Visual Thinking" has become required reading in fields ranging from data visualization to cognitive psychology since its publication. Praised by visualization pioneer Edward Tufte and adopted by tech giants like Google and Microsoft for their design teams, the book revolutionizes how we understand the relationship between vision and cognition. Far from being a passive recording system, our visual perception is revealed as an active, selective process that shapes how we think, create, and interact with information.
Chapitre 2
The Visual Query: How We Actually See the World
The paradox of vision is that while we feel we see everything around us in complete detail, scientific evidence proves otherwise. We actually perceive remarkably little at any given moment-just what's needed for our current task. The solution to this paradox, as psychologist Kevin O'Regan puts it, is that "The world is its own memory." We don't need to keep a complete mental copy of our surroundings because we can simply move our eyes to sample any part of the environment within a tenth of a second.
This understanding fundamentally changes how we think about perception. Visual thinking isn't about passively receiving images but about actively allocating attention-from rapid eye movements to pattern recognition to the assignment of working memory resources. These acts of attention, which Ware calls "visual queries," drive our interaction with the world around us.
When we look at information displays like maps, graphs, or interfaces, we solve cognitive problems through a series of targeted visual searches for specific patterns. This process involves two complementary neural pathways: bottom-up processing (driven by visual information from the retina) and top-down processing (driven by our goals and attention). You can experience this duality by looking at an image containing both letters and faces-focusing on either makes the other recede from consciousness.
The bottom-up pathway processes information through three main stages. First, feature processing occurs in the primary visual cortex (V1), where approximately five billion neurons simultaneously analyze input from one million optic nerve fibers. Different types of specialized neurons detect edges, color differences, motion, and depth across the entire visual field-essentially functioning as "brain pixels."
At the intermediate level, these features combine to form patterns-visual space divides into regions of texture and color, while chains of features connect to form contours. Finally, at the highest level, processed information condenses into visual objects held in visual working memory. This system can only maintain about three objects simultaneously, which explains why people might fail to notice when speaking to a different person during the direction experiment.
While bottom-up processing moves from retinal image to features to patterns to objects, top-down processes operate in reverse at every stage. These processes, collectively called attention, are driven by our goals-whether physical actions like grasping a cup or cognitive tasks like understanding a diagram. Top-down attention biases processing in favor of what we're seeking. If looking for red spots, red-detecting neurons signal more strongly.
Chapitre 3
Designing for the Visual Brain
Understanding that we see through just-in-time visual queries has profound implications for information design. Effective displays must create environments where visual queries are processed rapidly and correctly for every cognitive task the display supports. This process is not unlike how a computer retrieves data from memory, but our visual system operates with remarkable flexibility and sophistication.
Consider the London Underground map-a brilliant example of design supporting specific visual queries. When planning a journey from Ealing to Clapham Common, users execute a series of visual operations: locating start and end stations, tracing colored lines, finding intersections, and estimating journey length. The map excels at supporting line tracing but sacrifices geographical accuracy, making it poor for estimating actual distances or finding nearby landmarks. This trade-off exemplifies how successful information design prioritizes specific user tasks over complete accuracy. Other transit maps worldwide have adopted similar principles, from New York's subway map to Tokyo's complex rail network visualization.
Our brains break these tasks into nested loops, operating simultaneously at different levels of abstraction. The outer loop handles general strategy-finding a map, identifying endpoints, planning operations-much like an executive control system. The middle loop executes visual searches to find patterns addressing the visual query, such as tracing colored contours representing train lines. The inner loop activates at each fixation point, rapidly evaluating patterns in the central visual field at about twenty per second. These loops work in concert, allowing us to seamlessly navigate complex visual information.
Unlike rigid computer algorithms, these nested processes are highly flexible and adaptive, drawing on experience-based visual search patterns we've developed for countless situations. For instance, when searching a crowded restaurant for a friend, we automatically filter by height, clothing color, and familiar movement patterns. The brain doesn't literally ask and answer questions but produces neural signals that change the status of objects in visual working memory, creating a dynamic interaction between perception and memory.
This distributed cognitive structure extends beyond our heads into what researchers call the "extended mind." Information stored in books, pictures, or computers functions similarly to information in memory, creating a seamless cognitive ecosystem. As cognitive scientist Don Norman noted, "The power of the unaided mind is highly overrated"-our real cognitive power comes from devising external aids that complement our visual thinking capabilities. This principle explains why tools like spreadsheets, diagrams, and visual analytics have become indispensable for complex problem-solving and decision-making.
Modern interface design increasingly recognizes these principles, creating displays that work with, rather than against, our visual processing systems. From smartphone interfaces to data visualizations, successful designs align with our natural visual query patterns while minimizing cognitive load.
Chapitre 4
The Pop-Out Effect: What Captures Our Attention
Visual attention operates like a flashlight beam sweeping through a dimly lit room-we move our focus systematically, point by point, picking out details in our environment. While we often navigate with only vague information to guide each eye movement, certain visual elements can be detected even in our peripheral vision, allowing us to rapidly and efficiently shift our attention. This selective attention mechanism helps us process the vast amount of visual information we encounter every second.
Some visual elements instantly "pop out" from a page, capturing our attention without conscious effort or serial searching. Pioneering research by Ann Triesman in the 1980s revealed that for certain targets, response time remained constant regardless of the number of distractors-suggesting parallel processing occurs automatically in the brain's visual cortex. The strongest pop-out effects occur when a single target differs from identical surroundings in basic features like color (red among green), orientation (horizontal among vertical), size (large among small), motion (moving among static), or stereoscopic depth (near among far). This phenomenon helps explain why we can instantly spot a yellow tennis ball in green grass or a moving animal against a still background.
These differences must be substantial to trigger the pop-out effect - for example, a thirty-degree orientation difference between lines, or highly contrasting colors. The brain processes these distinctions in under a tenth of a second, while non-pop-out patterns require multiple eye fixations taking several seconds to locate. Importantly, complex patterns like numbers, letters, or faces don't pop out regardless of how familiar they are-the features that create pop-out effects are hardwired into our visual system through evolution, not learned through experience.
To make elements easily findable in interface design, designers should differentiate them using primary visual channels. When multiple searchable elements are needed, different channels should be employed-form (orientation and size), color, and motion work as semi-independent processing channels. For example, in a complex data visualization, important data points might be marked by both color and size differences. When designing complex displays with multiple searchable symbols, using differences across multiple channels makes elements even more distinct. However, creating more than 8-10 independently searchable symbols is likely impossible due to the limited number of available channels and our cognitive capacity.
Motion proves exceptionally powerful for capturing attention, especially in our visual periphery. While our sensitivity to static detail drops rapidly away from the fovea (central vision), motion sensitivity remains strong throughout our entire field of vision-an evolutionary adaptation that helped our ancestors detect approaching predators or prey. Motion generates an almost irresistible "orienting response," drawing our attention automatically, though we quickly become habituated to continuous movement. For maximum effectiveness in interface design, signaling icons should periodically emerge and disappear rather than move continuously, preventing habituation while maintaining their attention-grabbing power. This principle is commonly applied in notification systems, warning signals, and interactive displays.
Chapitre 5
Organizing Visual Space: Patterns and Perception
Though we live in a three-dimensional world, human perception treats dimensions unequally. The two dimensions in the "picture plane" (up-down and sideways) differ fundamentally from the third dimension (towards-away). Each retinal location receives one color point from the world, with millions of points available across the up-down and sideways dimensions, but only indirect distance information in the away dimension.
This creates what's sometimes called "2.5D space" (though "2.05D" would be more accurate). We can sample new up-down and sideways information with rapid eye movements in under a tenth of a second, while sampling new depth information requires walking to a new location, taking seconds or minutes. This efficiency difference makes image-plane pattern processing far more developed in the brain than depth processing.
The binding process combines different features to identify contours or regions. Individual neurons in the primary visual cortex respond to oriented edge information and form mutual reinforcement networks with similarly oriented neighbors. These neurons fire in unison when stimulated by continuous edges while suppressing responses to differently oriented fragments. Top-down attention further enhances relevant edges while suppressing irrelevant ones.
Objects can be distinguished from backgrounds through various visual discontinuities including luminance, color, texture, and motion boundaries. The brain's generalized contour extraction mechanism processes these diverse inputs to discern object boundaries regardless of how they're defined. This explains why simple line drawings effectively convey complex objects despite bearing little physical resemblance to actual edges-they directly stimulate this generalized contour mechanism.
Visual interference occurs when similar features overlap-like interferes with like. Text becomes difficult to read when placed on backgrounds containing similar feature elements, even with different colors. To minimize interference, maximize feature-level differences between information patterns. Movement differences provide the most effective separation between overlapping patterns.
The up-down and sideways directions are special, perceived differently from each other and from other orientations. We're highly sensitive to whether something is exactly vertical or horizontal, judging this far more accurately than other orientations. This sensitivity stems from maintaining posture relative to gravity and from our modern world containing many vertical rectangles, which tunes more pattern-sensitive neurons to these orientations.
Chapitre 6
The Power of Color: Beyond Aesthetics
Color vision evolved to help us see objects like fruit, which is why fruit-eating animals have developed better color perception. While grazing animals and predators like cats have limited color vision, humans and other great apes have three dimensions of color perception.
In the retina, three types of cone cells-short, middle, and long-wavelength sensitive-make color vision fundamentally three-dimensional. Images of human retinas reveal far fewer short-wavelength-sensitive cones (blue), which are also less sensitive to light. This explains why small blue text on dark backgrounds or yellow text on white backgrounds is difficult to read-there's insufficient luminance contrast for clear visibility.
In area V1 of the visual cortex, cone signals are transformed into three color-opponent channels: red-green, yellow-blue, and black-white (luminance). The red-green channel represents the difference between middle and long-wavelength cone signals, allowing us to detect subtle red-green contrasts despite overlapping cone sensitivities. The luminance channel combines outputs from long and middle-wavelength cones, while the yellow-blue channel represents the difference between luminance and blue cone signals.
The most fundamental principle for using color in design is that detailed information requires luminance contrast. Black on white provides maximum contrast, but excellent results can also be achieved with yellow on black or dark blue on white. For small text, the ISO recommends a luminance ratio of at least 3:1 between text and background, severely limiting color options to dark text on light backgrounds or vice versa.
Color coding is primarily used to indicate categories of information, as seen in land-use maps. When designing color codes, visual distinctness (supporting visual search) and learnability (colors that "stand for" specific entities) are paramount considerations. The unique hues-red, green, yellow, and blue-should be used first, followed by consistently named colors like pink, brown, orange, grey, and purple.
There are strict limits to how many colors can be used effectively as codes, with studies recommending between 6-12 colors for complete reliability. This limitation exists because background colors can distort the appearance of small symbol colors, causing confusion. In complex designs with both color-coded symbols and background regions, small areas should be strongly colored with black-white channel contrast against larger background areas.
Chapitre 7
Visual Space and Information Access
Moving our eyes is our lowest-cost method for gathering environmental information, requiring so little cognitive effort that we're unaware of making several eye movements per second. In contrast, physically searching through disorganized filing cabinets or traveling long distances for meetings represents a much higher cost of information acquisition that can disrupt our train of thought.
While we live in a three-dimensional physical world, our perceptual experience is more accurately described as 2.5-dimensional. We have rich information about the "up" and "sideways" dimensions directly from retinal images, but much less information about the "towards-away" dimension, which comes from depth cues.
These cues can be divided into pictorial cues (reproducible in photographs) and non-pictorial cues. Pictorial cues include occlusion, perspective effects, shadows, height on picture plane, shading, depth of focus, size relative to known objects, and atmospheric contrast. Each depth cue has unique properties that can support different kinds of visual queries and can be applied independently according to design goals.
The concept of affordances, developed by psychologist James Gibson in the 1960s, revolutionized perception theory by claiming we perceive physical possibilities for actions rather than just retinal images. A flat ground surface affords walking, a horizontal surface at waist height affords support for objects, and certain objects afford use as tools. Perception of space fundamentally concerns action potential within our environment.
Different information access methods have different cognitive costs: internal pattern comparison takes about 0.04 seconds, eye movements 0.1 seconds, mouse hover queries 1.0 second, and mouse selections 1.5 seconds. High-performance computer graphics have enabled non-metaphoric navigation methods like zooming interfaces, where scale changes exponentially, allowing movement from an Earth overview to a molecule in under 30 seconds.
The functional aesthetics of space relate directly to the cost of accessing information. Though we live in three dimensions, we always view the world from a particular viewpoint, and redirecting our gaze is always faster than moving our heads or bodies. For interactive applications, designers must consider the costs of information access, comparing eye movements to mouse clicks, or virtual walking to flying or zooming.
Chapitre 8
Visual Objects and Meaning
The inferotemporal cortex contains neurons specialized for complex visual patterns corresponding to recognizable objects and scenes. It includes subareas like the fusiform gyrus (responding to faces) and regions specialized for cars and houses, with links to areas providing multimodal connections between visual and other sensory information.
Object recognition relies on several sets of neurons responding to a few "generalized views" (like full face, three-quarters, profile), with each providing a distinct visual pattern. Because component pattern recognizers tolerate distortion, the overall recognition mechanism does too. This isn't matching to idealized forms but neural responses to ranges of patterns.
People can identify scene categories (busy road, rural landscape, fast-food restaurant) in less than a tenth of a second, even for unfamiliar specific scenes. This rapid "gist" perception happens as quickly as object identification, challenging theories that scenes are identified through their component objects. Scenes have characteristic spatial feature components-distributed patterns of textures and colors-that enable recognition without object identification.
Objects and scenes gain meaning through links to information stored in specialized brain regions, including action patterns, eye-movement scanning strategies, and semantic content in language systems. These links activate through "working memories"-temporary groupings that form between active visual patterns and non-visual stored meanings, particularly information in verbal working memory.
The brain has specialized neural subsystems for language processing (including Wernicke's area for interpretation and Broca's area for speech production) distinct from visual processing areas. Verbal working memory holds about two seconds of speech information in an "echoic loop"-approximately three chunks of information. Much of what we consider "thinking" is internalized speech in this verbal working memory.
Visual working memory can hold between one and three objects depending on their complexity, creating a significant bottleneck in visual thinking. This capacity critically influences design effectiveness-when thinking with graphic images, we constantly gather information chunks, hold them in working memory, formulate queries, and relate them to new information. This explains why side-by-side comparisons are vastly more efficient than page-switching, as eye movements are ten times faster than changing pages.
Chapitre 9
Bridging Visual and Verbal Thinking
The chapter challenges the cliche "A picture is worth a thousand words," noting that some concepts (like conditional statements about halibut prices) resist visual expression. Good design isn't about choosing between pictures and words but understanding when each is most effective and how they should be combined.
Sign languages provide valuable insights into the nature of language and visual expression. Contrary to common belief, sign languages weren't invented by hearing people as translations of spoken languages-they emerged spontaneously when deaf children gathered in schools, developing their own grammar and vocabularies distinct from spoken languages.
Language fundamentally consists of socially developed shared symbols with grammar, expressible through speech, writing, or signing. These symbols are arbitrary-the word "dog" has nothing inherently dog-like about it. Visual thinking, by contrast, relies on pattern perception that comes partly from evolution and partly from visual experience-not from social convention.
Natural language incorporates a form of logic distinct from visual representation, using qualifiers like "if," "and," "but," and "otherwise" that enable abstract reasoning and conditional instructions. Visual representations can incorporate logic, but it's the logic of pattern, object, and space rather than abstract verbal logic.
Even before speech development, infants communicate by pointing. Deictic gestures-actions that provide the subject or object of speech by directing attention to objects-remain crucial throughout life. Beyond finger pointing, gaze direction and body orientation can also serve as deictic references, with gestures varying in precision to express certainty or uncertainty about what's being indicated.
The purpose of narrative is to capture the audience's cognitive thread-the sequence of concepts held actively in visual and verbal working memories. In successful presentations, audience members' cognitive threads roughly follow the author's designed narrative thread. Unlike information seeking (which is internally driven), narrative presentations guide viewers to look at the same visual objects in the same order, activating similar concepts.
Chapitre 10
Creative Visual Thinking: The Design Process
Creative visual thinking doesn't happen solely in the mind. While ideas may originate there, the major work of creative design occurs through dialogue with rapid production media-sketches for artists, clay models for sculptors, diagrams for engineers and scientists. This externalization process follows common patterns across disciplines, beginning with concept formation, followed by externalization through rough sketches, constructive critique where the design is visually tested, and finally consolidation where the design is modified and refined.
Lines are visual chameleons, capable of representing many different things because they trigger our edge-detection mechanisms. A simple circle can have five different meanings (coin, ball, ring, hole, or plugged hole). The power of lines lies in how they activate our generalized contour extraction processes. In Massironi's scribble exercise, random looping lines become birds with minimal additions, demonstrating how our perception constructs meaning from minimal cues.
Unlike sketches that serve as visual prototypes, diagrams express concept structures that may not be inherently spatial. They combine words with graphical elements to plan, design, and structure ideas across disciplines from architecture to science. While specialized fields have their own diagrammatic symbols, arrows deserve special attention as versatile abstractions that apply visual pattern-finding to abstract relationships.
Suwa and Tversky's study of architecture students revealed that designers don't simply transfer ideas from mind to paper-sketching itself is constructive. Designers create loose sketches, then interpret what they've drawn, seeing potential meanings emerge. This "constructive perception" is fundamental to design. As Bryan Lawson noted of architect Santiago Calatrava, his drawing isn't about producing artwork but understanding problems.
The choice between mental imagery and physical sketching follows cognitive economics. Visual working memory allows rapid concept manipulation but has limited capacity and permanence. Rough sketches take longer to create but support greater complexity and can be preserved or discarded. Detailed designs have the least flexibility but allow for elaboration. As design progresses, artifacts typically evolve from many low-cost sketches to fewer high-cost finished drawings.
Design thinking through sketching doesn't require conventional drawing ability-almost any scribble will do. The power comes from the visual interpretive skills of the creator. The effectiveness of sketching as a thinking tool stems from four factors: the line's ability to represent many things due to our visual system's interpretive flexibility, the speed of sketching and discarding ideas, the cognitive skill of interpreting lines differently and projecting new ideas onto partial scribbles, and the ability to mentally image additions.