AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: moe
此标签下有11条笔记。
2026年7月16日
Nemotron-Labs-Audex-30B-A3B (Audex): Unified Audio Intelligence Without Regressing on Text Intelligence
nvidia
audex
nemotron-cascade-2
audio-llm
moe
mamba-transformer
x-codec2
af-whisper
tts
text-to-audio
speech-to-speech
unified-audio-text
cascade-rl
cfg
2026年6月25日
M6: A Chinese Multimodal Pretrainer
multimodal
chinese
pretraining
moe
text-to-image
vqgan
encoder-decoder
mixture-of-experts
2026年6月25日
UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild
controllable-generation
controlnet
diffusion
multi-task
hypernet
moe
c2i
zero-shot
2026年6月25日
HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
text-to-image
dit
mmdit
moe
sparse-moe
flow-matching
distillation
dmd
gan
open-source
2026年6月25日
HunyuanImage 3.0 Technical Report
t2i
unified-multimodal
autoregressive
moe
diffusion
transfusion
chain-of-thought
open-source
rlhf
2026年6月25日
ICEdit:用 In-Context 生成范式做指令图像编辑(Enabling Instructional Image Editing with In-Context Generation in Large-Scale DiT)
image-editing
instruction-editing
in-context
flux-fill
lora
moe
inference-time-scaling
dit
rectified-flow
neurips2025
2026年6月25日
Ming-Omni / Ming-Lite-Omni: A Unified Multimodal Model for Perception and Generation
omni
moe
unified
any-to-any
speech
image-generation
video
gpt-4o-class
open-source
2026年6月25日
Wan 2.2
video-generation
t2v
i2v
ti2v
moe
diffusion-transformer
flow-matching
vae
open-source
cinematic
2026年6月25日
Qwen3.5-Omni Technical Report
omni
audio
audio-visual
speech
asr
tts
moe
thinker-talker
streaming
agent
multilingual
2026年6月25日
Seed3D 2.0: Advancing High-Fidelity Simulation-Ready 3D Content Generation
3d-generation
image-to-3d
pbr
vecset
rectified-flow
dit
moe
vae
articulation
scene-generation
simulation-ready
bytedance
2026年6月18日
新一代开源大 MoE 训练配方深挖(多方)
moe
llm
training-recipe
kimi-k2
minimax
hunyuan
skywork
step
dots-llm1
ling
ring
pangu
deep-dive