Hearing and vision are integrated continuously: the brain binds a voice to moving lips, a thump to the ball hitting the floor, provided they land inside a temporal window of very roughly a hundred milliseconds. Inside the window, vision can even drag sound’s apparent location, which is why cinema speakers can sit beside the screen while dialogue seems to come from mouths.
The binding window
The window is not symmetric. Because light outruns sound in the world, the brain tolerates vision leading audio far better than audio leading vision: a clap seen slightly before it is heard feels natural, the reverse feels broken almost immediately. The window also flexes with content. Speech binds tightly, simple flashes and beeps more loosely, and repeated exposure to a fixed delay recalibrates perception within minutes, which is why a slightly laggy system feels worse when you first sit down than it does an hour in.
Illusions at the seams
The McGurk effect shows vision editing what is heard: watching lips say “ga” while hearing “ba” produces a heard “da.” The double-flash illusion runs the other way, a single flash paired with two beeps is seen as two flashes. The ventriloquist’s act is the oldest demonstration in the trade: vision captures the voice’s location because the eyes report position more reliably than the ears.
Why it matters for visuals
Live visuals sit entirely inside this machinery. Reactivity reads as connection only when latency stays within the binding window, which is why frame-late visuals feel dead even to audiences who cannot say what is wrong. The flip side is generous: inside the window, the brain actively glues picture to sound and will credit the visuals with precision they do not literally have. Land the kick within a frame or two and perception does the rest of the work for free.