Sound, the Interaction Signature the Web Has Left Unused
Since 2018, browsers have blocked every sound until the visitor's first gesture. Sound nonetheless remains the least occupied interface channel on the web. This article explains what it adds, when to use it and how not to annoy.

Since 1 July 2021, a new electric car sold in the European Union is no longer allowed to be silent. Regulation 540/2014 (opens in a new tab) requires an acoustic alerting system, active up to 20 km/h and when reversing, with a minimum level of 56 dB at 20 km/h, according to the European Blind Union (opens in a new tab), which campaigned for the measure. A silent object deprived pedestrians of information that a combustion engine had always given without anyone thinking about it. The legislator had the noise put back.
The web went the other way. Between June 2017 and March 2019, Safari, Chrome and then Firefox decided that no site could emit a sound before the visitor had clicked or touched the page. That rule ended background music and videos that started on their own. It also installed a broader convention, that of a website which makes no noise. This convention is not a property of the medium. Operating systems, game consoles, payment terminals and native apps use sound to confirm, to warn and to sign. Interface sound design, meaning the choice of sounds a product plays in response to actions, is an established discipline there. The web is one of the few interactive environments where silence is the norm. The channel is therefore almost unoccupied.
How the web became silent
The silence of the web is a decision taken by three browsers in response to specific abuses. On 8 June 2017, WebKit, the engine behind Safari, announced (opens in a new tab) that Safari on macOS High Sierra would block autoplay of media with sound by default on most websites, using an inference engine that decides site by site. The advice to developers was to assume that any playback of an audio or video element requires a click. In April 2018, Chrome 66 applied its own autoplay policy (opens in a new tab) to audio and video elements. Google says it blocks roughly half of unwanted autoplays. Chrome 71, in December 2018, extended it to the Web Audio API, the programming interface used to generate and process sound inside a page. Firefox 66 followed (opens in a new tab) in March 2019.
| Date | Browser | Rule |
|---|---|---|
| June 2017 | Safari, macOS High Sierra | Media with sound no longer start on their own on most sites. An inference engine decides site by site |
| April 2018 | Chrome 66 | Audible audio and video elements require a gesture from the visitor, except on sites where they already consume media |
| December 2018 | Chrome 71 | The Web Audio API starts in a suspended state until the visitor has interacted with the page |
| March 2019 | Firefox 66 | Audible audio and video content is blocked by default |
The decisions that made the web silent by default
Chrome grants one measured exception. Its media engagement index calculates, site by site, whether the visitor regularly consumes media there for more than seven seconds with sound on. A video or music site thus earns the right to start on its own. An interface sound, short by nature, never meets that condition. The W3C Web Audio specification (opens in a new tab) then wrote the rule into the standard. An audio context is created in a suspended state and the browser may refuse to move it to a running state until the page has received user activation, meaning a click, a tap or a key press.
The second lock is in the visitor's pocket. In 2014, a study by Martin Pielot and colleagues, presented at the MobileHCI conference, followed 15 Android users for a week. They received 63.5 notifications a day on average (opens in a new tab) and only 46.2% of them arrived on a phone in normal mode. The rest landed on a device set to vibrate (41.5%) or to silent (12.2%). In 2016, publishers interviewed by Digiday estimated that up to 85% of the views (opens in a new tab) of their videos on Facebook happened with the sound off. Facebook never published that figure, it came from the publishers themselves. It was still enough to impose burned-in captions on every social video.
These two locks produced a design convention. Since sound cannot start before a gesture and will often be cut on arrival, teams stopped designing any. Chrome's usage counter gives an indirect measure of that absence. On 23 September 2026, a Web Audio context was being created on 5.1% of page loads (opens in a new tab), up from 1.3% at the start of 2018. That counter measures the creation of a context, not the playing of a sound. A good share of those creations serves to identify devices for advertising purposes. The order of magnitude still speaks. Sound is absent from the vast majority of pages and that absence is neither technical nor regulatory. It comes from a design habit.
What sound adds that a pixel does not
Sound arrives before the image. A 2015 study by Jain and colleagues on 120 medical students (opens in a new tab) measured a significantly shorter reaction time to an auditory stimulus than to a visual one, in men and women alike. An older measurement on 14 subjects gave 284 milliseconds on average for sound (opens in a new tab) against 331 for light. The gap amounts to a few tens of milliseconds. On an interface, it does not make the action faster. It makes the confirmation perceptible before the eye has even checked the screen.
The second contribution is independence from the gaze. William Gaver, a researcher at Apple in the 1980s, was the first to formulate it. In a 1986 paper he proposed auditory icons (opens in a new tab), caricatures of natural sounds that inform about the source of an event, like crumpling paper for a discarded file. In 1989 he demonstrated them on the Macintosh with the SonicFinder (opens in a new tab), arguing that sound should be used in a computer as it is used in the world, where it conveys information about the nature of the event that produces it. The same year, Meera Blattner and colleagues defined the rival family, earcons (opens in a new tab), abstract sound motifs of one or more notes, assembled into families to represent related actions. Almost every sound in a modern operating system descends from these two families.
The effect has been measured. In 2002, Stephen Brewster of the University of Glasgow published a series of experiments (opens in a new tab) on a Palm III, the personal organiser of the time, whose screen was tiny. Participants entered five-digit codes on buttons 16 pixels wide and then 4 pixels wide, in silence or with a confirmation sound on every press. With the enriched sound, the gap in codes entered between large silent buttons and small sonified buttons fell to three over fourteen minutes. Both audio treatments significantly lowered the workload measured by the NASA-TLX questionnaire and no significant effect on annoyance was found. In a second experiment run while walking outdoors, participants preferred the sonified buttons. The author's conclusion is that a small button with sound becomes as usable as a large silent one.
One clarification is needed about what the literature contains. That study remains, to date, the strongest quantified case on interface sound. The often-cited 2008 experiments on touchscreen keyboards, by Eve Hoggan and Stephen Brewster, deal with tactile feedback (opens in a new tab), meaning vibration. They say nothing about sound. No published A/B test shows that an interface sound changed a website's conversion rate. Sound is therefore not justified by a promise of conversion. It is justified by what it does to perception.
On that point the research is clear. In 2004, Massimiliano Zampini and Charles Spence at Oxford had participants bite into crisps while modifying, in real time, the sound of their own chewing. When the volume was raised or only the high frequencies, between 2 and 20 kHz, were amplified, the crisps were judged crisper and fresher (opens in a new tab), although nothing about them had changed. Sound modifies the judgement made about the material of an object. That is the mechanism the sound of a car door or a switch exploits and it is the one a silent site gives up. The visitor there judges the material of a button on its appearance alone.
Sound, a brand asset almost nobody occupies
In February 2020, Ipsos published a meta-analysis of 2,015 American video ads (opens in a new tab) drawn from its testing database. It ranks each ad by the attention it earns for the brand, in three tiers. It then looks at which brand assets appear in it. Visual assets, logo, colour or slogan, appear in 92% of cases. Audio assets, a sonic signature or music, appear in 8% of cases. Ads that use an audio asset are 3.44 times more likely to sit in the top tier. The ratio rises to 8.53 for a sonic signature. For visual assets the ratio is 1.15. The metric concerns the attention paid to a brand inside a video, not the memorability of an interface. Transposing the coefficients would be an abuse. What the study establishes is simpler. The audio channel is almost empty and, when it is occupied, it weighs more than the saturated visual channel.
Brands that have understood this take sound out of advertising. In February 2019, Mastercard introduced a sonic identity (opens in a new tab) designed as a complete architecture, from musical scores and ringtones to the acceptance sounds played by payment terminals. The press release quantifies no effect. It shows where a company puts sound when it takes it seriously, on the approved payment, an interaction that is rare, decisive and repeated over the months.
Apple goes further in its recommendations for visionOS, the system of its headset. Its interface guidelines state (opens in a new tab) that people generally keep sound on while wearing the device and that an app which plays no sound, especially in an immersive moment, can feel lifeless and may even seem broken. Silence, in an environment where sound is expected, reads as a failure. On the web it reads as the norm, which leaves a site free to choose. Layouts are converging (opens in a new tab), the same components and the same typefaces appear everywhere. A confirmation sound designed for one site appears nowhere else. The visitor who hears it has no other site to compare it with.
When interface sound design belongs on a website
Operating system guidelines have already settled the question of frequency. Google's Material guide ranks sounds (opens in a new tab) in four levels. Hero sounds mark a rare moment, a success or a major step. Primary UX sounds accompany frequent actions and must bear being heard often without feeling annoying or redundant. In 2024, Google's sound team summed up the rule (opens in a new tab) in one sentence. The more often an interaction happens, the less intrusive its sound should be. Overuse takes relief away from the moments meant to be highlighted.
| Level | Frequency | On a website |
|---|---|---|
| Hero | Rare | The order placed, the appointment confirmed, the end of an immersive experience |
| Primary | Frequent | A primary button, a form submission, a state toggle |
| Secondary | Occasional | An input error, a notification inside the interface |
| Ambient | Continuous and decorative | The sonic texture of a 3D scene or an immersive page, which the visitor can cut with one gesture |
The four sound levels of the Material guide and what they become on a website
The same guide devotes a section to silence, judged as important as sound. It rules out three situations. Interfaces that call for discretion, users who have asked not to be interrupted and actions performed very often do not need sound. Moving between pages, hovering a menu or scrolling fall into the third category. Apple expresses the same idea through the behaviour of the iPhone's silent switch. When it is flipped, only the sounds people explicitly initiate (opens in a new tab) keep playing, like a video they started or an alarm. Keyboard clicks, sound effects and interface feedback are cut. Pielot's study confirms it another way. Its data contains no evidence that a phone in silent mode slows the response to notifications. Sound therefore does not serve to trigger the action. It serves to qualify the action once it is done.
Accessibility sets the last limit. WCAG criterion 1.4.2 (opens in a new tab), at level A, requires that a sound which starts on its own and lasts more than three seconds can be paused or adjusted independently of the system volume. The reason is concrete. A person browsing with a screen reader hears the page through speech synthesis and a background sound covers it. An interface sound lasting a fraction of a second, triggered by a gesture, does not fall under that rule. An ambient background does. One gap remains that the web has not filled. CSS offers preference queries for reduced motion, contrast, transparency or data (opens in a new tab), which a site can read in order to adapt. None exists for sound and no known proposal goes in that direction. A site that plays sounds must therefore offer its own visible setting and remember it from one visit to the next.
What the browser imposes and what it leaves to choose
An audio context created before the visitor's first gesture is born suspended and the page has to wake it in the wake of a click or a tap, a sequence that Chrome documents (opens in a new tab). The first sound of a site is therefore, at the earliest, a response to the first gesture, never a greeting. That constraint suits interface sound, which by definition responds to an action.
A subtlety of the iPhone then decides what is heard. On Safari for iOS, sounds produced by the Web Audio API go through the ringer channel and the silent switch cuts them. A video or an audio file played by a page element goes through the media channel, which the switch does not touch. A report filed with WebKit in March 2022 describes this behaviour (opens in a new tab) and Apple's answer. A page's audio session is of the ambient type by default. It is therefore muted when the phone is on silent. Since iOS 17, a page can declare itself of the playback type to override that. This declaration is the subject of a W3C specification (opens in a new tab), published as a Working Draft on 13 November 2024. For an interface sound, the right choice is not to override. Apple's doctrine files interface feedback among the nonessential sounds that silent mode must cut. A site that bypasses that setting plays a sound the visitor has explicitly refused.
What remains is the craft of the sound itself. Apple's guide recommends slightly randomising the pitch and volume of a single file on each playback rather than multiplying files, which avoids the mechanical repetition of an identical sound. Blattner's earcon principle completes that rule. The sounds of one site form a family, in which the sound of an error and the sound of a success share a timbre and differ by motif. The visitor then recognises the voice of the site before recognising the message.
An empty channel on a web that looks alike
The silence of the web is a response to abuses, decided by three browsers between 2017 and 2019 before being written into the specification. It protects the visitor and it sets the frame. That frame forbids any sound before the first gesture, any sound that overrides silent mode and any background that covers a screen reader without a control to stop it. Within that frame, a site keeps the possibility of answering an action with a sound, on the rare moments that Google's and Apple's guides reserve for that use, an order placed, an appointment confirmed or the end of a journey.
Websites converge on the same layouts and the same components. The audio channel has stayed empty there out of habit rather than constraint. A site that places a short, rare signature there, one the visitor can cut with a gesture, occupies a dimension of the experience that almost every other site leaves unoccupied.