-
FlowRL: Matching Reward Distributions for LLM Reasoning
Paper • 2509.15207 • Published • 119 -
Kwaipilot/KAT-Dev-72B-Exp
Text Generation • 73B • Updated • 32 • • 156 -
Agentic Entropy-Balanced Policy Optimization
Paper • 2510.14545 • Published • 108 -
Multi-Agent Deep Research: Training Multi-Agent Systems with M-GRPO
Paper • 2511.13288 • Published • 19
Malkesh Dalia
malkesh2911
·
AI & ML interests
None yet
Recent Activity
liked a model 12 days ago
CohereLabs/command-a-plus-05-2026-w4a4 upvoted a paper 13 days ago
Lance: Unified Multimodal Modeling by Multi-Task Synergy upvoted a paper 13 days ago
TradingAgents: Multi-Agents LLM Financial Trading FrameworkOrganizations
None yet