The Rise of AI Image and Video Generation
Generative image and video tools have moved from novelty to production tool in a couple of years. Full article coming soon.
AI-generated images went from a curiosity that produced recognisably strange output to a genuinely production-ready tool in a remarkably short window. Video is following the same curve a few years behind, and the pattern of what changed — and what it means for people who make visual work for a living — is worth taking seriously rather than dismissing as a passing trend.
Fact: current generation image models produce output that's frequently indistinguishable from professional photography or illustration at a glance, with substantially better instruction-following and consistency than the models from just a couple of years earlier — fewer stray extra fingers, more reliable text rendering, better adherence to detailed, multi-part prompts. Video generation has moved from a few seconds of unstable, low-resolution motion to coherent clips with reasonable temporal consistency, though it still lags image generation in reliability and control.
What actually changed the trajectory
The jump from impressive-but-unreliable to usable-in-real-work came less from a single breakthrough and more from better controllability — the ability to specify composition, lighting, style and detail precisely enough that the output matches an actual creative brief rather than a rough approximation of one. That's the difference between a novelty generator and a production tool: not that the images look better in isolation, but that a professional can get a specific, intended result reliably enough to build a workflow around it.
Analysis: the honest read on displacement isn't "AI replaces visual creators," it's that the value shifts from execution speed to direction and taste, similar to the pattern in software development. Someone who already has a strong visual sense — knows what makes a composition work, what a brand needs, what a specific creative brief actually requires — gets dramatically more productive with these tools. Someone without that sense gets fast, mediocre output that still needs a trained eye to fix or reject. The tools amplify existing taste rather than substituting for it.
Opinion: the genuinely difficult, unresolved part of this shift isn't technical — it's the honest question of provenance and consent around training data and style, which the industry hasn't settled and probably won't settle cleanly for years. Using these tools well and using them responsibly are two different skills, and treating "the technique makes something possible" as the same thing as "the technique is fine to use" is a mistake worth naming directly rather than glossing over.
Prediction, held loosely: video generation likely follows image generation's trajectory from novelty to genuine production tool over the next few years, narrowing but probably not eliminating the reliability gap between the two. The more durable skill through all of it is likely to be the same one that mattered before generative tools existed — a trained eye for what actually looks and works well — applied to a much faster, cheaper production process rather than replaced by one.
Written by
Gehna Stavonin-de Montagnac
Writing on artificial intelligence, software, automation, business and finance.
Related articles
Creating Images with Nano Banana Pro: Notes from Actually Using It
Hands-on notes on generating images with Nano Banana and Nano Banana Pro, building a phone-friendly version of an AI studio, and the idea of an image's 'JSON DNA.'
ChatGPT vs Claude vs Gemini vs Grok: How the Major AI Platforms Differ
The major assistants are converging on capability and diverging on judgement, interface and who they're really built for. A practical comparison rather than a leaderboard.