Sik.limited Logo

Cafe Owners, Is That an AI Background Music Playlist You're Playing?

More than half of the music uploaded to streaming services every day is now AI-generated, and a lot of it ends up on cafe speakers as free four-hour YouTube playlists. This piece walks through why regulars notice within days, why the customers who notice first are the ones you can least afford to lose, and how little you actually save by playing it.

Sik ·

Last month I spent about two hours in a cafe in Seongsu, Seoul's design district. Concrete walls, lighting done right, coffee at 7,000 KRW (about $5) a cup. I found a seat and opened my laptop.

About thirty minutes in, the music started bothering me.

It had been fine at first. Some pleasant jazz-adjacent pop, volume set about right. But three songs in, then four, something snagged. The words in the lyrics kept repeating. Coffee, rain, jazz, Sunday. The voice changed from track to track, but what it had to say never did. Then one chorus landed on a line that amounted to "there's nothing I can do about the rainbow."

Who sings that, resigned, like it's a real problem? I opened Shazam.

No result.

I ran it again. Same thing. That's when it clicked. These weren't written, they were generated. The stuff you find all over YouTube under titles like "Cafe Vibes Pop Songs 4 Hours."

Half of it gets made. About 3% of it gets heard.

Let's skip the vibes and start with numbers.

Deezer, the streaming service, tracks and publishes the volume of new tracks uploaded each day. As of July 2026, more than half of daily uploads are AI-generated. That's roughly 90,000 tracks a day. A year and a half ago it was 10,000.

The same announcement carried another number alongside it. AI tracks account for 1–3% of actual plays, and 85% of even that is flagged as fraudulent and excluded from payouts.

Half of everything made, and almost nobody listening. That's what's coming out of the cafe speaker. Not music somebody chose to hear, but music left over precisely because nobody chose it.

And these places are noticeably multiplying. It's not just independent cafes. Select shops, mall food courts, the same sound.

It holds up for a few days. Then it doesn't.

Most customers won't catch it on a first visit. I didn't.

Pull one AI-generated track out on its own and it's pretty convincing. The mix is clean, the audio quality is fine. The problem is that store music doesn't stop at one track. You loop a four-hour playlist all day, and a customer hears forty minutes to two hours of it. A regular hears it two or three times a week.

That repetition is what gives it away. Whenever background music in Korean cafes comes up in English-language forums, the reaction is always some version of the same thing: didn't notice at first, but after a few days this singer just keeps saying the same stuff. Coffee, rain, jazz, Sunday. The unease I felt in thirty minutes turns into certainty for a daily regular within a few days.

It isn't a question of whether you get caught. It's a question of when. And the answer is: fast.

Ears don't have eyelids.

An ugly sign you can just not look at. Turn your head, close your eyes, done. Sound doesn't work that way. It comes into your ears the entire time you're sitting there.

There's one concept worth leaning on here: processing fluency. Rolf Reber, Norbert Schwarz and others formalized it, and the gist is simple. The easier a stimulus is to process, the more people like it. And usually they have no idea why they liked it. They misread the sensation of easy processing as "hey, this is nice."

Music is a stimulus where fluency does especially heavy lifting. You're predicting the entire time you listen. You track how a player pushes and pulls against the beat, you anticipate where the singer will breathe, you sketch out where the next bar is headed. When the prediction lands, processing load drops, and we feel that as comfort.

From here on this is my own hypothesis. Tracks made with generation tools keep missing those predictions. A phrase moves on before it resolves, the emotional line runs flat where it should rise, consonants smear slightly. The lyrics are grammatical but you can't quite get hold of what they mean. Like that narrator giving up on the rainbow.

When predictions keep failing, the load never comes down. The customer doesn't articulate "the lyrics to this song are weird." Instead, twenty or thirty minutes in, they get restless. They can't focus, their shoulders tighten, they want to get up for no reason they can name.

"For some reason it's hard to sit here for long."

The moment that sentence shows up, the shop has already lost the customer. And the customer doesn't know why either.

Counterargument: isn't slow music actually a good thing?

Here's where this piece takes its sharpest hit. Let's deal with it.

The classic in store-music research is Ronald E. Milliman's 1986 restaurant study in the Journal of Consumer Research. He swapped the background music in a working restaurant and measured customer behavior. With slow music, dwell time ran about 25% longer than with fast music, and the longer people stayed, the more drinks they ordered.

But the variable Milliman manipulated was tempo, and only tempo. Not whether the music was good or bad.

And what exactly is on those AI playlists in cafes? Jazz-adjacent, lo-fi, gentle piano loops. Almost all of it slow. Apply Milliman's conclusion straight across and AI lo-fi should be keeping people in their seats longer.

The objection is valid. So we have to separate the axes.

Tempo is a dial for arousal level. Fast and the body hurries, slow and the body slackens. Processing fluency is the condition under which that dial works at all. If the brain is continually failing to predict the sound, no amount of slowness will get anyone to relax.

It isn't that slow music relaxes people. It's that predictable slow music relaxes people. The tracks Milliman played were, of course, written by humans and performed by humans. Predictability wasn't even in the study as a variable. He didn't control for it; in 1986 it simply couldn't be a problem.

Now that condition is wobbling. So if you want to use Milliman as a counterargument, you have to restate the question: is slow enough, or does it have to be slow and predictable? I think it's the latter. But nobody has measured that question yet, so I'll leave it as a hypothesis rather than a conclusion.

How long they sit decides how much they spend.

Set the tempo debate aside, because Milliman did establish one thing firmly. Dwell time and spending are attached. It was measured in a real business forty years ago and it hasn't budged since.

A cafe doesn't sell coffee. It sells a seat and a stretch of time. The second transaction happens when a customer is comfortable enough to stay a while.

  • The conversation runs long, so they grab a slice of cake from the case
  • The laptop session keeps going, so they order another americano
  • On the way out they pick up a bag of beans or some drip packs

Shorten dwell time and all three disappear at once. The customer drinks their drink and leaves.

Isn't faster turnover a good thing? That's a conversation for shops with a line out the door. In a place with no queue, faster turnover just means empty seats sooner.

What you save versus what you're betting.

From here on it's assumption. No study has measured how many minutes of dwell time AI music costs you. Pretending otherwise would collapse this piece along with it. So instead of asserting a loss, let's calculate the threshold — how little you'd have to lose before the savings stop being savings.

Start with the savings side. For a domestic store-music service, a coffee shop of 50–99㎡ runs about 23,000 KRW (about $17) a month. Public performance royalty is billed separately, and it's smaller than people expect: 4,000 KRW (about $3) a month for the same size bracket, and under 50㎡ there's no obligation at all. Call it 27,000 KRW (about $20) a month all in.

So what's the threshold?

In terms of a 6,500 KRW (about $5) dessert, it's four of them a month. Not four a day. Four a month. Sell one fewer every other day and you're out 100,000 KRW (about $74) a month; one fewer every day and it's 195,000 KRW (about $145).

Measured in regulars, the threshold drops further. One party that used to come twice a week stops showing up, and roughly 50,000 KRW (about $37) a month vanishes. One party. Twice what you saved.

I don't know the probability. I don't and nobody does. But when the threshold sits this low, the bet is strange even without knowing the odds. The savings side is fixed at about 20,000-something KRW a month, while the losing side starts at a single party of customers.

The first customers to go are the ones you can least afford to lose.

Here's the real problem: who notices the wrongness first.

The oblivious customer never figures it out. "Just some pop music," and on with their day. But that customer is typically one americano and thirty minutes.

The ones who catch it are a different group. The freelancer who comes two or three times a week and works for three or four hours. The person who uses the same cafe for every meeting. The one who buys beans. The people actually holding up your revenue. They're sensitive to sound environments — because it's their job, or their taste, or at minimum because they sit there a long time.

And these customers say nothing. They don't write it in a review either. Who docks a star over weird music? They just start sitting in the cafe two blocks over the following week.

This doesn't only happen with sound. Not long ago a cafe in San Francisco put AI-generated food photos on its menu and took them all down after customer backlash. Someone said the croissant looked like a Lego block.

But the two cases shouldn't be lumped together. A menu photo is a representation of the product. If the thing I'm buying doesn't match the picture, that's deception, and the grounds for anger are clear. Background music isn't that. Nobody comes in to buy the music.

Only one thing carries over. Customers can tell the difference between what a shop prepared and what a shop just filled in. The difference is that with a photo they can point a finger and complain, while with music they simply stop coming.

Five minutes will tell you what you're playing.

That's the diagnosis. Here's what you can actually do.

One, identify it. Listen to three consecutive tracks off your current playlist. Two or more of the four below and it's close to certain.

  • Shazam or SoundHound returns no result
  • The voice changes track to track but the lyrical subjects overlap (coffee, rain, seasons, days of the week)
  • The lyrics are grammatical but you can't get hold of what they mean
  • The channel keeps posting four-hour videos every few days

Two, replace it. Three options, different in character.

  • Store music services — around 20,000-something KRW a month for a coffee shop. Some products cover the public performance royalty on your behalf or bundle it outright. Simplest option, and you never think about copyright
  • Commercially licensed music libraries — for when you want to pick yourself. More work, but the shop's own character comes through
  • AI music services curated for commercial use — a different thing entirely from what you scrape off YouTube. I've written up how to judge these separately in Is there an AI playlist that's actually okay for a cafe?

Three, operate it. Don't run the day as one block. Instrumentals during the morning work hours, vocals allowed in the afternoon conversation hours as long as the lyrics don't jump out, tempo nudged up in the last hour before close. Three files is enough.

Four, volume and speaker placement. If a speaker sits directly above a seat, no music on earth will keep a customer there long. The standard is simple: music should be slightly buried under the sound of people talking. If the music beats the conversation, it's already noise.

What's coming out of your speakers right now?

A good cafe has different air the moment you walk in. Music sits naturally in the gaps between the grinder and the steam wand. The customer relaxes without ever consciously registering it. They're not sitting there to admire the music; they're sitting there longer because of it.

When something's off, the opposite happens. The customer is uncomfortable without being able to name the cause. And since they can't name it, they never ask for it to be fixed. They just don't come back.

In case this reads wrong, I'm not saying using AI-made music in a shop is inherently a mistake. The problem isn't the tool, it's that nobody chose. 90,000 tracks a day pour out, and you take four random hours of it and press play. That choice is the signal that reaches your customer.

I left half a cup of coffee in Seongsu that day. The coffee was good. That's the part that bothered me.

Latest posts