← All topics

Tag

Diffusion Transformer (DiT)

Every Diffusion Transformer (DiT) story we've curated in Bowl of Data, newest issue first — part of our weekly digest across AI, security, blockchain, and engineering.

3 items · 3 issues

Beats AI & ML
Often covered with LoRA 1 KV cache 1

Week 38 · 2026

Read the issue →
  • VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention

    VC-Attention is a new low-bit attention kernel that optimizes video diffusion transformer inference by addressing value quantization errors and softmax bottlenecks. It utilizes value smoothing via k-means clustering and a fused probability casting method to achieve high fidelity and significant hardware acceleration.

    AI & ML HuggingFace Papers Source ↗

Week 30 · 2026

Read the issue →
  • Self Gradient Forcing: Native Long Video Extrapolation

    The researchers present Self Gradient Forcing (SGF) to solve the lack of gradient flow in historical KV caches during autoregressive video generation. This method enables much more stable and consistent long-form video extrapolation without the massive memory overhead of full backpropagation.

    AI & ML HuggingFace Papers Source ↗

Week 28 · 2026

Read the issue →