Samsung Research and Samsung's Mobile eXperience division detailed the technology behind the latest Galaxy Z Fold8 generation's Portrait Video feature on September 16.
The feature originally debuted alongside Samsung's new foldables at Galaxy Unpacked in July. The latest technical explanation focuses on what never appears cleanly in a specification table: how a smartphone recreates an optical effect its tiny camera cannot naturally produce to the same degree.
The basic concept remains familiar. The main subject stays sharp while the background becomes soft.
The important change is that the blur is no longer intended to behave like one uniform filter placed behind a segmentation mask.
A real lens does not blur every part of the background equally
With a large-aperture lens, blur depends partly on the distance between the focus plane and the other objects in a scene.
Something just behind the subject may remain reasonably recognizable, while a much more distant background can become heavily defocused.
Small bright sources also turn into bokeh highlights whose size, shape and intensity depend on the optics and depth.
Samsung says recreating that appearance therefore requires much more than applying a Gaussian blur to everything behind a person.
Its researchers first analyzed how light spreads through actual large-aperture camera lenses and converted the resulting bokeh characteristics into data the video-processing system could use.
AI then provides an approximate map of distance
The second part of the system comes from depth estimation.
The phone uses AI to infer the relative distance of different elements inside the scene.
That information is combined with the optical model to determine how much blur should be applied and how large bokeh highlights should become.
The intended result is a gradual transition between depth planes rather than a binary division between “sharp subject” and “blurred background.”
That gradual behavior is what helps synthetic shallow depth of field resemble an optical one.
Hair remains the natural enemy of synthetic bokeh
Samsung immediately identifies one of computational photography's oldest problems: boundaries.
AI-estimated depth alone cannot always determine exactly where a person or object ends.
Hair, thin branches, glasses, fingers and partially transparent objects can easily be assigned to the wrong plane.
The familiar failure mode is a portrait where strands of hair disappear into the blur while small pieces of background between those strands remain unnaturally sharp.
Samsung says it therefore combines estimated depth with edge information extracted from the original video.
The system is effectively asking not only “how far away is this region?” but also “is there a credible visual boundary here?”
One bad frame is annoying; video has to remain stable across every frame
Video makes the problem significantly harder than still photography.
A segmentation error in a photograph remains in one place.
In video, the same error can change from frame to frame, producing an edge that flickers, shimmers or jumps around the subject.
Temporal consistency therefore matters almost as much as the accuracy of each individual frame.
Samsung says teams repeatedly recorded footage across changing weather, lighting, subject movement and distances to identify unstable cases.
Software optimization and additional AI training data were then used to reduce rendering that still appeared artificial.
The phone can even shift focus naturally to another person
Portrait Video recognizes faces and can move the focus to another person as the main subject turns away or the composition changes.
The behavior is intended to approximate a decision a camera operator or autofocus system might make on a conventional camera.
The phone then has to modify both subject separation and the distribution of blur around the new focus plane.
A sudden binary switch would immediately expose the effect as synthetic.
Samsung therefore aims for a more natural transition in depth and bokeh.
Focus can also be changed after recording
Computational rendering has one major advantage over physical optics.
After capture, the user can select a different subject to emphasize.
Blur strength can also be changed on a scale from 0 to 7.
The phone therefore retains enough scene information to recalculate part of the depth separation after recording.
That is fundamentally different from a traditional large-aperture lens. On a conventional camera, missed focus during a shot is usually baked into the footage.
Here, part of the optical appearance remains a software decision that can be changed afterward.
Samsung is even modelling how bright points spread out of focus
Bokeh does not merely mean that an area is blurry.
It also describes how out-of-focus elements are rendered, particularly small sources of light.
A streetlamp or string light behind a subject can become a much larger luminous disk through a wide-aperture lens.
Samsung says it measured those characteristics from real optics in order to model the effect more naturally.
The system then adjusts the size of those highlights according to estimated scene depth.
That is a meaningful step beyond early portrait modes that tended to turn the entire background into one uniformly softened surface.
All of this still has to run while the phone records 4K video
The hardest constraint may simply be compute time.
Samsung says the processing must operate on 4K video in real time.
For every frame, the system needs enough analysis to estimate depth and edges, track subjects and reconstruct optical rendering without falling behind the incoming video stream.
Video quickly turns a technique that may be acceptable for one photograph into a thermal-management problem.
Running the full pipeline dozens of times per second increases power draw and heat, particularly inside something as compact as a foldable phone.
Samsung deliberately removed calculations users were unlikely to notice
The research team says it had to prioritize where compute resources were spent.
Processing judged essential to image quality was preserved, while calculations with little visible impact were reduced.
That is an important reality hidden behind the phrase “on-device AI.”
A model can theoretically produce a more sophisticated result with more compute. On a phone, it also has to finish before the next frame and remain efficient enough not to drain the battery or overwhelm the thermal envelope.
The best mobile algorithm is therefore not necessarily the one that performs the most operations.
Synthetic depth and optical depth are still fundamentally different
Samsung describes rendering inspired by professional lenses, not a physical transformation of the camera module.
A real wide-aperture lens creates shallow depth of field as light passes through the optical system and reaches the sensor.
The smartphone first records an image with its own optical properties, then estimates scene structure and reconstructs a different appearance.
That leaves computational failure modes such as incorrect masking, bad depth estimation or unnatural rendering around complicated geometry.
The advantage works in the opposite direction: the effect can be adjusted after capture and changed in ways a fixed physical lens cannot retroactively reproduce.
The Fold8 is becoming a camera with partially software-defined optics
Computational photography has existed on smartphones for years, but Portrait Video shows how its role is shifting.
Software is no longer limited to correcting noise, exposure or color produced by the sensor.
It is trying to reconstruct an entire optical characteristic: depth of field, transitions between planes and bokeh rendering.
Part of the “lens choice” therefore becomes processing.
The same physical camera can produce different interpretations of depth without physically changing aperture or optics.
Samsung's own shooting advice reveals where the illusion works best
The company recommends backgrounds containing small bright light sources and enough depth separation for the bokeh to become visible.
Backlighting or side lighting can also make hair and subject contours easier to distinguish.
Those recommendations sound remarkably similar to advice for shooting with an actual fast lens.
That is revealing. Even when blur is generated computationally, a scene providing useful depth cues and suitable highlights remains easier to render convincingly.
Software can simulate optics, but it cannot make every composition equally interesting.
The next smartphone-camera improvement may be increasingly difficult to put on a specification sheet
Previous generations were easy to market through sensor size, megapixels and zoom ratios.
None of those numbers really explains this particular change.
The improvement comes from a depth model, boundary analysis, measurements derived from physical lenses and a processing pipeline efficient enough to run continuously.
That is considerably harder to reduce to one headline specification than a 200-megapixel sensor.
It also reflects a deeper trend: smartphone rendering depends less exclusively on the optics physically attached to the phone and increasingly on the optics its software is capable of reconstructing.