🚀ReVisual-R1 is a 7B open-source multimodal language model that follows a three-stage curriculum—cold-start pre-training, multimodal reinforcement.
Shawn
csfufu
AI & ML interests
None yet
Recent Activity
authored a paper about 18 hours ago
Unify-Agent: A Unified Multimodal Agent for World-Grounded Image Synthesis authored a paper about 19 hours ago
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence? upvoted a paper 2 days ago
Agentic-MME: What Agentic Capability Really Brings to Multimodal Intelligence?