How We Made a Cinematic AI Music Video for Under $1,000
A four-minute music video with two invented civilisations, a cast of named characters, a city built from nothing, and a Dolby Atmos mix. No location. No crew. No set build. Total production cost: under $1,000.
That number is the headline, so let me put it up front and then spend the rest of this piece on the part that actually matters — because "we used AI, so it was cheap" explains almost nothing about why this video looks the way it does. Plenty of people have the same tools and the same thousand dollars. The difference is in where we refused to let the model make the decisions.
When Girinandh came to me with the idea of presenting GAIA in a very mysterious way, I nodded like I knew exactly what he meant.
I did not know what he meant. I sat with it for a while afterwards, genuinely confused, turning the word over. Mysterious. What does mysterious look like? Because "mysterious" is not a shot list. It isn't a colour palette or a lens choice or a location. It's a feeling, and feelings are the hardest brief in the world to receive, because two people can agree completely on the word and be imagining two entirely different films.
Hi — this is Manish, and this is the making of the music video for GAIA.
I want to write this one out properly, because the process behind it was genuinely unusual, and because most "how we made it with AI" posts skip the part where you actually have to make a thousand decisions that have nothing to do with the tools. The tools were the easy part. They always are.

- The whole city was modelled in Blender first, so every camera angle was a decision — not a prompt outcome.
- Character consistency was solved structurally, with locked reference elements, not with better prompting.
- Finished with a Topaz upscale pass so the delivery held up at full resolution.
- Under $1,000 all-in — the saving came from replacing a set build, not from cutting craft.
Decoding a one-word brief
Here's what I've learned about abstract briefs: the worst thing you can do is ask the client to be more specific. Not because it's rude, but because it usually doesn't work. If Girinandh could have described the video precisely, he would have. "Mysterious" wasn't laziness — it was the most accurate word available to him for something he could feel in the music but hadn't yet seen.
So instead of asking him to define it, I went to the track. The composition is doing something specific: it moves between registers, it holds back, it withholds resolution longer than a pop arrangement would. Ranina Reddy's vocal sits on top of it not as a lead line demanding attention but as something older, more like an invocation. The world drums underneath it — Ramana and Bharath — give it a foundation that isn't from any one place you can point to on a map.
Listening to it repeatedly, "mysterious" started to resolve into something I could actually build toward. The music wasn't asking for fog and shadows. It was asking for scale you can't account for. The feeling you get from a civilisation whose ruins are too large and too precise for the story you were told about who built them. Mystery not as darkness, but as unexplained magnitude.
That reframing was the whole unlock. Once "mysterious" became "build something enormous and don't explain it," I had a job I could brief a team on.
The worst thing you can do with an abstract brief is ask the client to be more specific. Go to the work instead. The answer is usually already in it.
Two regions, one myth
What we landed on was two worlds. Not two locations — two civilisations, with different rulers, different architecture, different relationships to the element they lived in.
The first is Araka. The second is Neervi.
Araka is a surface kingdom, and its ruler is its queen, Ara. Neervi is submerged, and its ruler is Gamini. GAIA moves between them, and the tension between the two is the spine of the whole piece.
Now — if Araka sounds familiar to anyone who watched the BOI BOI music video, that's not a coincidence, and I'd been waiting a long time for someone to notice. Araka is the same exact location we showed in BOI BOI. Same world. Same architecture.
GAIA is, effectively, the prequel to BOI BOI. Where BOI BOI showed you Araka as a place that simply existed, GAIA shows you its birth.
I can't overstate how much that decision changed the project. The moment we committed to GAIA being a prequel rather than a standalone, every design question got a second constraint: not just "does this look right," but "does this credibly become the thing we already showed people?" That's a harder problem, and a much more interesting one.
Ara, and the bridge between two universes
My team and I sat through hours — genuinely hours, across multiple sessions — of character design, trying to work out how to fit the BOI BOI universe into the GAIA universe without either one feeling retrofitted.
This is the part of the work that never photographs well. There's no impressive screenshot of four people arguing about whether a crown reads as "founding monarch" or "late-period decadence." But it's where the video was actually made.
The problem was structural. BOI BOI had its own visual grammar, established and already public. GAIA needed its own identity — it couldn't just be BOI BOI with different music, or the prequel reveal would land as a rerun rather than a revelation. But it also couldn't drift so far that the connection felt like a marketing claim rather than something visible on screen.
Somehow, the character of Ara turned out to be the bridge between the two worlds.
Once we had her — once her design was locked — everything else oriented itself around her. She carries the visual DNA that BOI BOI viewers would recognise, while sitting at the origin point of a story they'd only seen the later chapters of. She's the hinge the two universes turn on. Design her right and the connection is self-evident; design her wrong and no amount of edit rhythm would have sold it.
I'd argue this is the single most transferable lesson from the whole project, and it has nothing to do with AI: when you're connecting two worlds, find the one element that belongs to both, and solve that first. Everything downstream gets easier.
The court around her
Ara doesn't appear alone, and that turned out to be a bigger job than it sounds. A ruler with no one around her doesn't read as a ruler — she reads as a woman standing in a landscape. So the Order was designed too: the attendants who move with her, dance in the bathhouse sequence, and raise the monuments in the finale.
Five of them, each given a name and a distinct silhouette — Pyra, Zephra, Sela, Naia and Lyra. They had to be recognisably a group, sharing a visual language of veils, robes and coin-metal ornament, while still being individually identifiable when the camera picks one out of a crowd. That balance is the whole trick with an ensemble: same world, different people. Too uniform and they're wallpaper; too varied and the group stops reading as an order at all.
Building a city in Blender so the camera could be real
Here's where the process gets genuinely interesting, and where I think we did something worth other people stealing.
The obvious way to make an AI music video is to write prompts, generate shots, and pick the ones that look good. That workflow produces images. It does not produce coverage — a set of shots that feel like they were captured in one continuous space by one physical camera, which is the thing that separates a film from a slideshow.
So we didn't start with prompts. We started with architecture.
We built the city of Araka in Blender. An actual three-dimensional model of the place — its structures, its scale, its spatial relationships. Not for final render. Not for beauty. For navigation.
Because once the city existed as geometry, I could move a camera through it freely. I could fly down a processional route and find the angle where the towers stack correctly against the horizon. I could drop to ground level and see what the scale actually feels like from a person's eyeline. I could circle a structure until the composition clicked — the same way you'd walk a real location on a recce, except the location was one we'd invented.
And when I found a composition I liked, I took that exact frame — the real one, from the real geometry, with the real perspective — and brought it into Higgsfield, and prompted precisely what I wanted that frame to become.
We didn't prompt our way to a city. We built the city, moved a camera through it like a location scout, and then prompted the frames we'd already chosen.
Blender — the blockout
The same frame, generatedThat inversion is everything. In the standard workflow, the model decides the composition and you accept or reject it. In ours, we decided the composition — with full spatial authority, based on a space that genuinely existed in three dimensions — and the model handled surface, light, material and detail.
The difference shows up in ways that are hard to name but easy to feel. Sightlines are consistent. When you see a structure from two angles, it's the same structure. The camera moves like a camera with weight and intention rather than like a series of unrelated viewpoints. The city has a geography you could, in principle, draw a map of — because we did draw a map of it, first, in Blender.
Two specifics from that build are worth stealing, because both cost us time before they saved us any.
The rock formations are photogrammetry, not procedural noise. Our landscape reference was Egypt's White Desert — a flat chalk plain studded with wind-eroded formations, not rolling dunes. We tried generating those shapes procedurally, repeatedly, and every attempt produced smooth blobs that read as CG the instant you looked at them. What finally worked was taking real photogrammetry rock scans, scaling them far past their intended size, and keeping their native textures. Actual erosion has a logic that noise functions don't reproduce, and the eye knows it immediately.
The greybox renders carry a depth pass. That's the hinge the whole pipeline turns on: geometry plus depth going into the AI stage is what lets the model repaint a shot's surface while leaving its camera and its space untouched. Without depth you get a beautiful image that has quietly stopped agreeing with the shot before it.
We also specified lens character the way you would on a live shoot — Helios 44-2 58mm swirl for the intimate close-ups, 2x anamorphic for the hero wides — and locked the finish spec early and never moved it: 2.39:1, 35mm grain, halation, deep blacks. "Cinematic" is a set of specific optical behaviours, not a vibe, and it's much easier to hit when you write those behaviours down before you generate anything.
It is more work upfront. Considerably more. You are building a set that will never be seen in the form you built it. But it front-loads every hard decision into a stage where changes are cheap, and it means that by the time you're generating, you're executing a plan rather than gambling.
Neervi: what if the Chola empire was underwater?
Neervi came from a single question, and I still think it's the best question anyone asked during this project: what if the Chola empire was underwater?
Not a generic sunken city. Not Atlantis with Indian ornamentation bolted on. Specifically the Chola empire — one of the longest-running and most architecturally sophisticated dynasties in South Indian history — reimagined as a civilisation that developed beneath water rather than beside it.
That specificity did an enormous amount of work. Because the moment you anchor to a real historical culture, you inherit an entire design language: the temple forms, the metalwork, the proportional systems, the ornamental vocabulary, the way authority is expressed through structure. You're no longer inventing from nothing. You're asking a much more productive question — what would this have looked like, under these conditions?
And underwater changes everything downstream. Weight behaves differently. Light behaves differently — it comes from above, diffused, in shafts rather than in a single hard source. Fabric doesn't fall, it drifts. Scale reads differently because the medium itself is visible. Architecture that would need buttressing on land doesn't, and structures that would be impossible in air become plausible.
Gamini rules this world. The name comes from kami — the Japanese word for a god or divine spirit — which tells you how she was meant to be read: not a monarch who happens to live underwater, but something the people of Neervi would have understood as divine. She had to feel like she belonged to that world, not like a surface character dropped into water for a sequence.
The Gamini problem: history vs. what you can actually show
This is the part of the process I most want to write down, because it's a real constraint that real productions hit and almost nobody discusses honestly.
When you research the actual Chola empire — as opposed to how it's depicted in Hollywood or Bollywood — you find that women were, historically, mostly topless. That's simply what the record shows. The costuming you see in mainstream period films is a modern invention, applied backwards onto a culture that didn't dress that way.
We obviously cannot show that in a music video.
So we had a genuine design problem, and it wasn't a trivial one. Ignore the history and dress Gamini in film-convention Chola costume, and we're doing exactly the thing we set out not to do — repeating a cinematic cliché instead of engaging with the actual culture. But historical accuracy was off the table for reasons that need no defending.
The other thing we ruled out early was the easy escape hatch: put her in a traditional saree and move on. It would have been instantly readable and completely wrong — a saree drapes and settles in a way that belongs to a different era and, more to the point, to a different medium. It's the costume shorthand every period production reaches for, and reaching for it would have told the audience they were watching a costume drama rather than a civilisation.
What followed was a lot of trial and error. Genuinely a lot. I went through iteration after iteration trying to arrive at an outfit that met three conditions at once: it had to be recognisably rooted in Chola design language, it had to read as regal enough for a ruler, and it had to be completely non-explicit.
The version we landed on takes the ornamental vocabulary of the period — the metalwork, the layering logic, the way status is signalled through material rather than coverage — and rebuilds it as something that stands on its own. It's not a historical reconstruction and it isn't pretending to be. It's a design that acknowledges where it came from without either sanitising the source or reproducing it literally.
I'd rather be honest that this was a compromise than pretend it was a pure creative choice. Most costume decisions in most productions are compromises. The craft is in making the compromise look inevitable.
There's a second, entirely unromantic reason this had to be solved properly rather than fudged: AI video tools refuse explicit nudity outright. Every generation that drifted too far in that direction simply failed. So the constraint we'd have applied for editorial reasons anyway was also being enforced by the pipeline — which, in practice, meant the design had to be genuinely good rather than merely covered. You cannot prompt your way around a refusal. You have to actually design the costume.
The consistency problem — the thing that actually kills AI video
If you've tried to make anything longer than about fifteen seconds with generative video, you already know the real enemy. It isn't image quality. Image quality has been good enough for a while now.
It's consistency.
A character whose face shifts between shots. Architecture that grows or loses a tier depending on the angle. A costume that changes its metalwork between the wide and the close-up. Individually these are small. Cumulatively they're fatal, because they break the one thing a narrative piece depends on: the viewer's belief that they're looking at a continuous world rather than a sequence of pictures.
Our answer was to stop treating each shot as an independent generation. To maintain consistency of every character across the entire video, we built elements in Higgsfield — locked references that every subsequent shot was generated against, rather than re-described each time.
That single decision, more than any other, is what held the video together across its full runtime. Not a better prompt. Not a better model. A structural choice to define the characters and the world once, properly, and then generate everything against those definitions.
In practice that meant every character got a full turnaround before they were allowed anywhere near a shot. Not one hero image — a set: front, side, back, the face uncovered, the eyes in close-up, the hands, the detail below the neck, and the costume in motion. Eight views, same light, same backdrop, every time.
The close-ups matter more than the wides here, which is counter-intuitive. A model will happily keep a silhouette consistent and then quietly redesign the hands, the jewellery or the weave of a veil the moment you push in — and close-ups are exactly where an audience looks hardest. Generating the detail views up front, and locking them, is what stops a character subtly becoming a different person every time the camera moves closer.
For the motion itself we worked with Seedance and Seedance 2.5. For character mapping we used GPT Image 2, generating the reference imagery before anything moved — establishing exactly who these people were as stills, locking faces, silhouettes and costume detail while they were still cheap to change, and only then putting them in motion.
And at the very end, everything went through a Topaz upscale pass. This matters more than it sounds. Generative video models output at resolutions that look fine in a browser preview and fall apart the moment the video is played full-screen on a decent display — soft edges, mushy detail in exactly the ornamental metalwork we'd spent weeks designing. Upscaling last, after the edit was locked, is what makes the difference between a video that survives a 4K playback and one that only ever looked good on a phone.
And here's the thing that surprised me most about the whole project: once our characters and locations were locked, making the video was pretty simple.
That sentence undersells weeks of work, but it's true, and it's the real lesson. All the difficulty in this production lived in the preparation — the Blender build, the character design sessions, the costume iterations, the reference elements. By the time we were actually generating shots and cutting them together, the hard problems were already solved. The video largely assembled itself, because we'd removed every decision that could have gone wrong at that stage.
All the difficulty lived in the preparation. By the time we were generating shots, there were no decisions left that could go badly wrong.
What it actually cost — and what that number hides
Under $1,000, start to finish. Here is what that figure is actually made of, and more usefully, what it isn't.
Blender is free. The city, the palace, the monuments, the walls, every camera move through Araka — all of it was built in open-source software running on a single desktop GPU. The largest creative component of this production had a software cost of zero.
Almost the entire spend is generation credits, across the image and video models, plus a Topaz licence for the final upscale. That's essentially the whole bill.
What we didn't pay for is where the number really comes from. No location fee. No set construction. No crew day rates, no travel, no equipment rental, no costume fabrication, no VFX vendor. A physical build of even one of these two worlds — a Chola-era palace, above water and below it — is a number with two more digits on it, and that's before anyone switches a camera on.
So the honest framing isn't "AI made this cheap." It's narrower and more interesting than that: AI collapsed the single most expensive line item in period fantasy production, which is building the world, and left every other cost roughly where it was. The song still needed real players — world drums, bass, a vocalist. The mix still went to a studio in Chennai and came back in Dolby Atmos. Those are real costs paid to real people, and no model replaces them.
One caveat I'd insist on, because budget posts always skip it: the time cost did not go down. It moved. Weeks went into the Blender build, the character design sessions, the costume iterations and the reference elements — work that happened before a single final frame was generated. If you're reading this hoping to make something like it in a weekend, the thousand dollars is achievable and the weekend is not.
AI didn't make the video cheap. It collapsed one line item — building the world — and left every other cost exactly where it was.
Why we're building a universe instead of videos
There's a strategic decision buried in all of this that I want to pull out, because it's the part I think other studios and other artists should copy.
Reusing Araka wasn't nostalgia. It was a deliberate choice to treat these videos as instalments in one continuous world rather than as self-contained pieces that happen to share an artist.
The economics of that are genuinely compelling. When we built Araka for BOI BOI, that design work was a cost absorbed by a single video. Building it again in Blender for GAIA — extending it, showing its origin, filling in what the first video only implied — meant that cost started amortising across two pieces instead of one. The character design, the architectural language, the ornamental rules we'd already established: none of it had to be re-litigated. We could spend our thinking on what was new.
But the real return isn't financial, it's attentional. A standalone music video competes for four minutes of a stranger's time and then it's over. A video that's secretly a prequel does something else — it retroactively adds meaning to something the audience already watched. Someone who saw BOI BOI a year ago and now sees Araka's founding has been given a reason to go back and rewatch with new information. That's not a view. That's a relationship.
And it changes what a viewer does with the next one. Once an audience understands that these pieces connect, they start watching for connections. They read the architecture. They notice when a motif recurs. They speculate in the comments about which world came first and who Ara really was. Attention you've earned that way is a completely different quality of attention than attention you've bought with a good hook.
I'll be honest about the cost, though, because there is one. Continuity is a commitment you can't easily undo. Every future piece set in this world inherits every decision we've already made — the architecture is fixed now, the timeline is fixed, Ara's design is fixed. We've traded away some creative freedom in exchange for accumulated meaning. If we'd made a mistake in the BOI BOI design language, we'd be living with it here, and we'd be living with it in whatever comes next.
That's a trade I'd make every time. Most creative work is disposable by default — made, posted, forgotten, replaced by the next thing four days later. Building something that accumulates instead of evaporates requires accepting constraints from your past self. The alternative is a body of work where nothing means anything beyond the week it launched.
A standalone video competes for four minutes. A prequel retroactively adds meaning to something they already watched. That's not a view — that's a relationship.
The objections a smart reader should raise
Let me argue against myself, because I would, if I were reading this.
"This is AI slop with extra steps." It's a fair thing to be suspicious of, given how much genuinely thoughtless AI video is being published right now. My answer is the Blender build. Nobody models a city they're never going to render in order to make slop. The entire pipeline we used exists specifically to take compositional authority away from the model and keep it with the director. You can disagree about whether the result is good — that's taste, and taste is allowed. But the choices in it were made by people, deliberately, over weeks.
"You're only using AI because you couldn't afford a real shoot." Partly true, and I won't pretend otherwise — no independent music video budget in India is building a Chola-era city, on land and underwater, with two rulers and full costume design. But "we couldn't have afforded it otherwise" and "this was the right tool" aren't mutually exclusive. There is no practical shoot that produces a submerged empire and a founding-era surface kingdom in the same 4:50. This is a class of image that didn't have a production route before.
"The historical stuff is a fig leaf." Reasonable challenge. The honest answer: the Chola grounding genuinely constrained the design — it's why Neervi has a coherent ornamental logic instead of generic underwater fantasy. But we deviated from the record where we had to, most obviously in Gamini's costume, and I've said so plainly above rather than claiming an authenticity we didn't earn.
"Does the prequel connection actually read, or is it just something you're claiming?" Best test I can offer: watch GAIA, then watch BOI BOI. Araka is the same place. If you don't see it, the connection failed and that's on us — but I'd rather you check than take my word for it.
What this project actually proves
I keep coming back to that first conversation, and how uncomfortable the word "mysterious" felt at the time.
Because the finished video isn't mysterious in the way I initially feared — dark, vague, hiding its weaknesses in shadow. It's mysterious in the way a real ruin is: entirely visible, rendered in detail, and still not fully explaining itself. That's a much harder thing to achieve, and you can only achieve it by building the world properly. You cannot suggest scale you haven't built. The audience can always tell.
Three things I'd want you to take from this if you're making anything in this space.
Build the space before you generate the shots. The Blender-first approach is more work and it's worth every hour. Compositional authority is the difference between a film and a gallery of images, and the only way to keep it is to have a real space to move a camera through.
Solve consistency structurally, not per-shot. Locked reference elements beat better prompting every single time. If you're re-describing a character in every generation, you've already lost — not because the model is bad, but because you've made continuity a matter of luck.
Constraints are the design. The Chola anchor gave Neervi its vocabulary. The BOI BOI continuity gave Araka its shape. Even the Gamini costume problem — which was purely a limitation — produced a more distinctive design than a free hand would have. Every genuinely interesting decision in this video came from something we couldn't do.
That's the part I'd push back on hardest in how AI production gets discussed. The tools remove the constraint of budget, and people treat that as the whole story. But removing every constraint doesn't give you better work — it gives you work with nothing to push against. The craft is in choosing which constraints to keep.
GAIA is out now on all official platforms. Stream it, and if you watch the video, watch the architecture.
Frequently asked questions
How much does it cost to make an AI music video?
GAIA cost under $1,000 in production, and that figure is almost entirely generative model credits plus a Topaz upscaling licence. Blender, which did the heaviest creative lifting, is free. The saving comes from eliminating location, set construction, crew and costume fabrication — not from spending less on craft. Music recording and the Dolby Atmos mix were separate, real costs paid to real musicians and a studio.
Which AI tools were used to make GAIA?
Blender for the 3D city build and camera work, GPT Image 2 for character mapping and reference stills, Higgsfield for locked reference elements and frame generation, Seedance and Seedance 2.5 for motion, and Topaz for the final upscale. Blender is the one doing the work most people assume the AI did.
How do you keep a character consistent across AI-generated shots?
Stop describing the character in every generation. Build locked reference elements once — face, costume, silhouette — and generate every subsequent shot against those references instead of re-prompting from scratch. Consistency is a structural problem, not a prompting problem, and no amount of prompt refinement fixes it.
Why model a city in Blender if AI can generate the image anyway?
Because prompting gives you images, not coverage. Modelling the city means you can move a camera through a real three-dimensional space, choose your composition the way a director would, and then generate from that exact frame. It keeps compositional authority with the filmmaker and makes sightlines agree from shot to shot.
Does AI video need to be upscaled?
In practice, yes. Generative video output looks fine in a browser preview and falls apart on a full-screen 4K display, particularly on fine ornamental detail. Run the upscale last, after the edit is locked, so you're not paying the processing cost on footage you later cut.
Credits
Music Composed & Arranged by C. Girinandh · Vocals Ranina Reddy · World Drums Ramana & Bharath · Bass Guitar Carl Fernandas · Dolby Atmos Mixed & Mastered by Bob Phukan at Aura Studios, Chennai · AI Visuals & Editing Manish & With Media.