简单来说,transformer是模型架构,AR/diffusion是利用这个架构进行训练/生成的方式
我看了下你上面提到的论文,2.2节第一句话就说了”LLaDA employs a Transformer [7] as the mask predictor, similar to existing LLMs.“
我看了下你上面提到的论文,2.2节第一句话就说了”LLaDA employs a Transformer [7] as the mask predictor, similar to existing LLMs.“


发表于 2026-1-19 18:20:41









































