Related Experiment Video
Updated: Jan 15, 2026

Photorealistic Learned Landscapes for Augmented Reality
Published on: June 27, 2025
HeadArtist-VL: Vision / Language Guided 3D Head Generation with Self Score Distillation
None:
We present HeadArtist-VL, a 3D head generation method that suits either vision or language input. With a landmark-guided ControlNet serving as a generative prior, we come up with an efficient pipeline that optimizes a parameterized 3D head model under the supervision of the prior distillation itself. We name such a process self-score distillation (SSD). In detail, given a sampled camera pose, we first render an image and its corresponding landmarks from the head model, and add some particular level of noise onto the image. When the input is a language prompt, we fed the noisy image, landmarks, and the language prompt into a frozen ControlNet twice for noise prediction. We conduct two predictions via the same ControlNet structure but with different classifier-free guidance (CFG) weights. The difference between these two predicted results directs how the rendered image can better match the language instructions. When the input is a reference image, we follow the aforementioned pipeline but with two modifications. First, we use an image encoder to obtain the image identity embedding, which is then sent to the ControlNet. Second, we use a novel-view diffusion model to synthesize the same reference image under the sampled camera pose to guide the self-score distillation process. In the experiments, our HeadArtist-VL produces high-quality 3D head sculptures with rich geometry and photo-realistic appearance, which significantly outperforms state-of-the-art methods. We also show that our method supports editing operations on the generated heads, including both geometry deformation and appearance change. 3D Head Generation and Editing, Vision / Language Guidance, Self Score Distillation.

