Patent issue dates cluster for administrative reasons, but a cohort granted to one company on a single day can still trace the outline of a product roadmap. On July 7, 2026, the US Patent and Trademark Office issued a set of patents assigned to Snap Inc., and read together they map the stages of an augmented-reality pipeline — seeing a scene, rendering into it, letting people interact, and presenting the result through display optics.
The anchor is US12675908B2, "Estimating 3D scene representations of images," naming inventor Titas Anciukevičius. The granted patent is directed to a machine-learning model that takes one or more ordinary 2D images of a real-world environment and produces a three-dimensional scene representation, separating the background from each object's own 3D shape, position, and appearance. Critically, the disclosure specifies that the model is trained in an unsupervised way — without manually labeled depth maps, segmentation masks, or object poses. The record describes a generative latent-variable design, per-object latents modeled as diagonal Gaussians, and a rendering network that maps points to color and density in a manner reminiscent of neural radiance fields.
The claimed scope is set out directly in the record:
A method comprising: receiving, by one or more processors, a set of two-dimensional (2D) images representing a real-world environment comprising a scene; sampling a three-dimensional (3D) scene representation from a posterior distribution conditioned on a single or multiple sets of images drawn from a distribution different than a training distribution; and generating, by a machine learning model, the 3D scene representation of the set of 2D images, where the 3D scene representation explicitly and separately defines a 3D shape and appearance of a background of the scene and a 3D position, 3D shape and appearance of each object of the scene depicted in the set of 2D images, where the machine learning model has been trained in an unsupervised approach from a dataset of images and their camera poses.— Estimating 3D scene representations of images, US12675908B2
Classified under CPC codes including G06T 7/75 and G06V 10/82, the grant is directed to pose estimation and neural-network image analysis. Its disclosure also reaches the payoff for a camera company: placing AR or VR virtual elements into video using the recovered 3D representation. That connects an abstract scene-understanding model to the practical business of putting convincing digital objects into footage from a phone or a pair of glasses.
The rest of the cohort fills in the pipeline
Where the hero grant handles perception, the same-day siblings handle what happens next. US12675840B2 is directed to late warping to minimize latency of moving objects, time-warping already-rendered content to an updated pose or object location so overlays stay locked as things move. US12675930B2 covers a state-space system for pseudorandom animation, describing a probabilistic motion model for animated content. Reconstruction, warping, and animation together describe the render half of an AR stack.
The interaction layer appears in two more grants. US12675169B2 is directed to gesture-based shared AR session creation, matching observed motion against a captured gesture to start a shared session between users. US12675157B2 covers pausing device operation based on facial movement — pausing AR sensors when a user's eyes look away, a power-and-attention detail specific to worn devices. And on the optics side, US12674986B2 is directed to an adjustable display arrangement for extended-reality devices, switching the virtual field-of-view mode of an XR display.
A grant records what a company has secured, not what it has committed to ship, and none of these documents names a product or a date. Still, the composition of the July 7 cohort is a signal in itself: it runs from unsupervised 3D perception through low-latency rendering and gesture interaction to the display hardware that shows the result. For a portfolio watcher, the pattern suggests Snap is protecting the full width of the AR pipeline rather than a single feature — with the reconstruction grant marking a specific bet that editable 3D scenes can be recovered from plain 2D images, no hand labeling required.
Comments
Loading comments…