預覽提示詞
Create a two-tier latent diffusion text-to-image schematic. Top row shows the forward diffusion process: a clean cat photo degrading left-to-right into pure Gaussian noise across T=1000 steps, labeled x0, x_t, x_T with the equation q(x_t|x_{t-1}) and a small beta noise-schedule curve rising from 0.0001 to 0.02. Bottom shows the reverse denoising U-Net: a symmetric encoder-decoder with skip connections, stacked ResNet plus self-attention blocks, a sinusoidal timestep embedding feeding each block, and cross-attention conditioned on a CLIP text-encoder of the prompt "a cat astronaut." Add a frozen VAE encoder/decoder bridging pixel and latent space. Include a 3x3 grid of generated sample thumbnails and an inset line chart of FID decreasing from 28 to 4.1 versus training steps (0 to 500k). Pastel panels, clean labeled arrows, publication-grade, calm and legible.




