AI Research 技术调研
Search
搜索
暗色模式
亮色模式
探索
标签: video
此标签下有25条笔记。
2026年7月16日
MC-JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Features
jepa
self-supervised
world-model
optical-flow
multi-task-learning
vicreg
convnext
video
motion-features
2026年7月16日
V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
world-model
jepa
self-supervised
video
action-conditioned
zero-shot-planning
robot-manipulation
predictive-embedding
non-generative
2026年6月25日
Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation
video
t2v
video-editing
one-shot
diffusion
fine-tuning
stable-diffusion
ddim-inversion
attention-inflation
2026年6月25日
Video Diffusion Models (VDM)
video
diffusion
3d-unet
factorized-attention
reconstruction-guidance
text-to-video
video-prediction
2026年6月25日
AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
video
t2v
motion-module
plug-and-play
diffusion
sd1.5
lora
motionlora
iclr2024
2026年6月25日
AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model
any-modality
multimodal-llm
llama-2
perceiver-resampler
frozen-llm
audio
video
imu
instruction-tuning
understanding
2026年6月25日
Pika 1.0
video
text-to-video
image-to-video
consumer-product
closed-source
generative-video
2026年6月25日
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
video
image-to-video
text-to-video
latent-diffusion
data-curation
open-source
edm
multi-view
Clips
2026年6月25日
VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
video
t2v
i2v
diffusion
lvdm
open-source
unet
2026年6月25日
VideoPoet: A Large Language Model for Zero-Shot Video Generation
video
autoregressive
llm
discrete-token
multimodal
text-to-video
image-to-video
audio
magvit-v2
zero-shot
2026年6月25日
Baichuan-Omni
omni
mllm
audio
video
image
speech
open-source
vita
siglip
whisper
conv-gmlp
2026年6月25日
Genie: Generative Interactive Environments
world-model
video
latent-action
unsupervised
maskgit
st-transformer
vq-vae
foundation-model
playable
agents
Params
2026年6月25日
HunyuanVideo: A Systematic Framework For Large Video Generative Models
video
t2v
dit
flow-matching
mmdit
3d-vae
open-source
scaling-law
2026年6月25日
LTX-Video: Realtime Video Latent Diffusion
video
t2v
i2v
dit
latent-diffusion
rectified-flow
video-vae
realtime
open-source
2026年6月25日
Open-Sora Plan / Open-Sora(2024 开源复现 Sora)
video
t2v
i2v
dit
diffusion
sora-reproduction
open-source
wf-vae
skiparse-attention
rectified-flow
2026年6月25日
MAGI-1: Autoregressive Video Generation at Scale
video
autoregressive
diffusion
flow-matching
chunk-wise
world-model
streaming
open-weights
2026年6月25日
Ming-Omni / Ming-Lite-Omni: A Unified Multimodal Model for Perception and Generation
omni
moe
unified
any-to-any
speech
image-generation
video
gpt-4o-class
open-source
2026年6月25日
Pika 2.0 / 2.1 / 2.2
video
t2v
i2v
keyframe
closed-source
consumer
pikaframes
pikadditions
2026年6月25日
Seedance 1.0: Exploring the Boundaries of Video Generation Models
video
t2v
i2v
dit
mmdit
flow-matching
rlhf
distillation
multi-shot
bytedance
doubao
jimeng
2026年6月25日
Show-o2: Improved Native Unified Multimodal Models
unified
understanding-generation
autoregressive
flow-matching
3d-causal-vae
video
qwen2.5
omni-attention
2026年6月25日
SkyReels-V2: Infinite-Length Film Generative Model
video
diffusion-forcing
autoregressive
flow-matching
dit
dpo
long-video
open-source
2026年6月25日
Sora 2
video
audio
text-to-video
world-model
diffusion
synced-audio
cameo
closed-source
2026年6月25日
VACE: All-in-One Video Creation and Editing
video
editing
unified
dit
controllable
inpainting
reference-to-video
wan
ltx-video
iccv2025
2026年6月25日
Wan: Open and Advanced Large-Scale Video Generative Models (Wan 2.1)
video
t2v
i2v
dit
flow-matching
3d-vae
wan-vae
open-source
video-editing
vace
2026年6月25日
Seedance 2.0
video
audio-video
multimodal
t2v
i2v
r2v
editing
binaural-audio
closed-source