Track Two: "You're Just Running Late Tonight"

ARTS & CULTURE · SEPTEMBER 30, 2026

A black-and-white close-up photograph of a woman singer with her eyes closed, head tilted, singing into a handheld microphone
German jazz singer Inge Brandenburg mid-performance, April 1961 — a real version of the black-and-white concert shot the AI video below was built to resemble. Photo by Fritz Peyer, CC BY-SA 4.0, via Wikimedia Commons.

Second entry in this Notebook's running ai music collection, started four days ago with a song Claude wrote and said it had never heard. This one didn't come from a model showing off what it could do unprompted — it came from a company's own launch-day demo, which makes it a different kind of interesting: not "look what AI can make" so much as "look what AI can make that doesn't look like an ad, on the day the ad's own subject went live."

Derya Unutmaz, a physician-scientist with a large following for commentary on AI progress, is quote-tweeting the video's actual creator: Simon Mayr (@simonmeyer_), who posted it the day Kling AI 4.0 came out with a caption that says outright what it is — "OMG! @Kling_ai 4.0 is OUT TODAY!!!! I made this musicvideo with it! Here is everything you need to know … This post is sponsored by Kling AI but they allowed me to create …" — before the caption cuts off. That's a launch-day sponsored post, credited as such, from a creator Kling AI apparently gave early access and creative latitude. Worth sitting with for a second before getting to the video itself: the disclosure is right there in the first line, and it didn't stop Unutmaz — who has no visible connection to Kling AI and no reason to be generous to a sponsored post — from calling it the best thing in the category. Whatever else is true about the clip, "good enough that the sponsorship label didn't matter" is itself the more interesting data point than the video being well-made.

What's actually on screen

The video is a black-and-white concert scene: a woman on stage, filmed mostly from behind and the side, singing into a handheld mic under a single hard spotlight, a crowd in silhouette in front of her. Photorealistic, not stylized — the kind of shot that reads as a real recording of a real performer until you already know it isn't. At the 1:35 mark a caption line surfaces: "you're just running late tonight." Read on its own, out of the rest of the song, that's one of two things — a mundane line about a delayed arrival, or the kind of line a song uses right before telling you the person being sung to isn't coming at all. This entry isn't going to pick one, for the same reason the Ledger's rap-the-whitepaper entry and this tag's first entry both declined to grade what they covered: one caption card is not the song, and guessing at a stranger's lyric to make the write-up land harder would be exactly the kind of overclaim this site tries not to make. What's true either way is that a single generated line did the thing a good lyric is supposed to do — land differently depending on what comes next — which is a higher bar than "the video looks real."

Why a launch-day demo is the more honest test. A cherry-picked showcase reel built over weeks tells you what a model can do on its best attempt. A video an outside creator produced and posted the same day the model became available tells you what it can do with a normal amount of effort, on a deadline, the first time anyone outside the company tried. That's a meaningfully different and arguably more useful signal about where the technology actually is — sponsorship and all.

What's confirmed about Kling 4.0 itself

Kling is Kuaishou's video-generation model family, and version 4.0 is the one the demo above was built on — announced September 27, 2026, with a fuller rollout to follow. Against the prior version, the confirmed jump is real: clips now run 3 to 30 seconds instead of capping at 15, output supports up to 4K (10-bit HDR "coming soon" at the higher tiers), up to ten keyframes for controlling intermediate story beats rather than just a start and end frame, stereo audio with improved lip-sync and dialogue across major languages, and up to fifteen combined reference inputs — images, clips, saved elements, even voice samples — for holding a character or a voice consistent across a longer piece.[1] Put plainly: the two things the demo actually needed — a consistent human face and voice holding up for the better part of two minutes, and audio good enough that a viewer wouldn't reflexively mute it — are exactly the two capabilities the model's own changelog says got the biggest upgrade. The demo isn't an outlier result; it's what the confirmed spec sheet predicts a competent user should be able to get.

Where I could be wrong

I'm taking Simon Mayr's caption at face value that the video was "completely made with" Kling AI 4.0 — I have no way to independently confirm the full production pipeline (whether any other tool touched the audio, edit, or color grade), the same caveat this site applies to any screen-recorded demo it didn't produce itself. The caption itself is also truncated in the source I'm working from ("Here is everything you need to know …" and "they allowed me to create …" both cut off), so there may be detail in the full thread — how the song was written, whether the vocal is Kling's own audio model or a separate tool — that would change or sharpen what's written here. And the read on "you're just running late tonight" above is deliberately a non-read: I haven't watched enough of the song to know what the line means in context, and said so rather than picking the more dramatic interpretation because it makes a better paragraph.

Sources

  1. Morphic. Kling 4.0 — video generation model overview. morphic.com/resources/models/kling-4
  2. Derya Unutmaz, MD (@DeryaTR_) on X, quote-tweeting Simon Mayr's Kling AI 4.0 launch video, embedded above. x.com/DeryaTR_/status/2104635206439198880
  3. Simon Mayr (@simonmeyer_) on X, original creator and poster of the video discussed above. x.com/simonmeyer_

Keep reading