Collect
previewQwen3-Coder: Agentic Coding in the World | QwenQwen3-Coder: Agentic Coding in the World | Qwenhttps://qwenlm.github.io/blog/qwen3-coder/ GITHUB HUGGING FACE MODELSCOPE DISCORD Today, we’re announcing Qwen3-Coder, our most agentic code model to date. Qwen3-Coder is available in multiple sizes, but we’re excited to introduce its most powerful variant first: Qwen3-Coder-480B-A35B-Instruct — a 480B-parameter Mixture-of-Experts model with 35B active parameters which supports the context length of 256K tokens natively and 1M tokens with extrapolation methods, offering exceptional performance in both coding and agentic tasks. Qwen3-Coder-480B-A35B-Instruct sets new state-of-the-art results among open models on Agentic Coding, Agentic Browser-Use, and Agentic Tool-Use, comparable to Claude Sonnet 4. 2026-07-24 05:44:06
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video GenerationStoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generationhttps://storydiffusion.github.io/ StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation 2026-07-24 05:44:06
DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid GuidanceDreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidancehttps://grisoon.github.io/DreamActor-M1/ DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance. 2026-07-24 05:44:06
Tongyi DeepResearch: A New Era of Open-Source AI Researchers | Tongyi DeepResearchTongyi DeepResearch: A New Era of Open-Source AI Researchers | Tongyi DeepResearchhttps://tongyi-agent.github.io/blog/introducing-tongyi-deep-research/ GITHUB HUGGINGFACE MODELSCOPE SHOWCASE From Chatbot to Autonomous Agent We are proud to present Tongyi DeepResearch, the first fully open‑source Web Agent to achieve performance on par with OpenAI’s DeepResearch across a comprehensive suite of benchmarks. Tongyi DeepResearch demonstrates state‑of‑the‑art results, scoring 32.9 on the academic reasoning task Humanity’s Last Exam (HLE), 43.4 on BrowseComp and 46.7 on BrowseComp‑ZH in extremely complex information‑seeking tasks, and achieving a score of 75 on the user‑centric xbench‑DeepSearch benchmark, systematically outperforming all existing proprietary and open‑source Deep Research agents. 2026-07-24 05:44:06
LeVoLeVohttps://levo-demo.github.io/ LeVo: High-Quality Song Generation with Multi-Preference Alignment 2026-07-24 05:44:06
previewLumiereLumierehttps://lumiere-video.github.io/ Space-Time Text-to-Video diffusion model by Google Research. 2026-07-24 05:44:06
previewOffer for youOffer for youhttps://forecast.mobilesl.com/eventsOur neural network makes predictions about what will happen in the world in the next 100 years.
Like
Kimi K2 ThinkingKimi K2 Thinkinghttps://moonshotai.github.io/Kimi-K2/thinking.html Kimi K2 Thinking, Moonshot's best open-source thinking model. 2026-07-24 05:44:06
X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent AttentionX-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attentionhttps://byteaigc.github.io/X-Portrait2/ X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention 2026-07-24 05:44:06
preview[2601.13976] FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation[2601.13976] FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigationhttps://arxiv.org/abs/2601.13976 Abstract page for arXiv paper 2601.13976: FantasyVLN: Unified Multimodal Chain-of-Thought Reasoning for Vision-Language Navigation 2026-07-16 14:27:27
preview[2601.14251] LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR[2601.14251] LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCRhttps://arxiv.org/abs/2601.14251 Abstract page for arXiv paper 2601.14251: LightOnOCR: A 1B End-to-End Multilingual Vision-Language Model for State-of-the-Art OCR 2026-07-16 14:27:27
preview[2601.14232] KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning[2601.14232] KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learninghttps://arxiv.org/abs/2601.14232 Abstract page for arXiv paper 2601.14232: KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning 2026-07-16 14:27:27
preview[2602.04804] OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models[2602.04804] OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Modelshttps://arxiv.org/abs/2602.04804 Abstract page for arXiv paper 2602.04804: OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models 2026-07-16 14:27:27
preview[2602.03510] Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers[2602.03510] Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformershttps://arxiv.org/abs/2602.03510 Abstract page for arXiv paper 2602.03510: Semantic Routing: Exploring Multi-Layer LLM Feature Weighting for Diffusion Transformers 2026-07-16 14:27:27
previewQwen3: Think Deeper, Act Faster | QwenQwen3: Think Deeper, Act Faster | Qwenhttps://qwenlm.github.io/blog/qwen3/ QWEN CHAT GitHub Hugging Face ModelScope Kaggle DEMO DISCORD Introduction Today, we are excited to announce the release of Qwen3, the latest addition to the Qwen family of large language models. Our flagship model, Qwen3-235B-A22B, achieves competitive results in benchmark evaluations of coding, math, general capabilities, etc., when compared to other top-tier models such as DeepSeek-R1, o1, o3-mini, Grok-3, and Gemini-2.5-Pro. Additionally, the small MoE model, Qwen3-30B-A3B, outcompetes QwQ-32B with 10 times of activated parameters, and even a tiny model like Qwen3-4B can rival the performance of Qwen2. 2026-07-24 05:44:06
previewDiffusionLight: Light Probes for Free by Painting a Chrome BallDiffusionLight: Light Probes for Free by Painting a Chrome Ballhttps://diffusionlight.github.io/ We present a simple yet effective technique to estimate lighting in a single input image. Our research uncovers a surprising relationship between the appearance of chrome balls and the initial diffusion noise map, which we utilize to consistently generate high-quality chrome balls. We further fine-tune an LDR diffusion model (Stable Diffusion XL) with LoRA, enabling it to perform exposure bracketing for HDR light estimation. Our method produces convincing light estimates across diverse settings and demonstrates superior generalization to in-the-wild scenarios. 2026-07-24 05:44:06
preview[2602.07120] Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model[2602.07120] Anchored Decoding: Provably Reducing Copyright Risk for Any Language Modelhttps://arxiv.org/abs/2602.07120 Abstract page for arXiv paper 2602.07120: Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model 2026-07-16 14:27:27
previewAI or HumanAI or Humanhttps://ai-or-human.github.io/ Play the game and try to guess the difference between what is AI or Human. 2026-07-24 05:44:06

0.0460 seconds

You can install the AI app from our store.

Scan QR code to get a link to APK file