Yinxu Pan

cppowboy

https://github.com/Cppowboy

AI & ML interests

RL for LLM, Code&Math Reasoning, Function Calling, Code Interpreter, Vision-Language Pretraining

Recent Activity

upvoted a paper 1 day ago

SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks

liked a dataset 4 days ago

mercor/APEX-SWE

liked a dataset 5 days ago

mercor/apex-agents

View all activity

Organizations

upvoted a paper 1 day ago

SlopCodeBench: Benchmarking How Coding Agents Degrade Over Long-Horizon Iterative Tasks

Paper • 2603.24755 • Published 3 days ago • 22

liked a dataset 4 days ago

mercor/APEX-SWE

Updated 4 days ago • 3.49k • 18

liked a dataset 5 days ago

mercor/apex-agents

Viewer • Updated 26 days ago • 480 • 40.3k • 104

upvoted a paper 5 days ago

LongCat-Flash-Prover: Advancing Native Formal Reasoning via Agentic Tool-Integrated Reinforcement Learning

Paper • 2603.21065 • Published 7 days ago • 74

New activity in Qwen/Qwen3.5-397B-A17B 5 days ago

Can not reproduce evaluation results on SWE-Verified

#63 opened 17 days ago by

cppowboy

upvoted a paper 9 days ago

MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild

Paper • 2603.17187 • Published 11 days ago • 132

upvoted 5 papers 10 days ago

Online Experiential Learning for Language Models

Paper • 2603.16856 • Published 11 days ago • 57

TRUST-SQL: Tool-Integrated Multi-Turn Reinforcement Learning for Text-to-SQL over Unknown Schemas

Paper • 2603.16448 • Published 11 days ago • 58

InCoder-32B: Code Foundation Model for Industrial Scenarios

Paper • 2603.16790 • Published 11 days ago • 301

MiroThinker-1.7 & H1: Towards Heavy-Duty Research Agents via Verification

Paper • 2603.15726 • Published 12 days ago • 181

SWE-Skills-Bench: Do Agent Skills Actually Help in Real-World Software Engineering?

Paper • 2603.15401 • Published 12 days ago • 18

New activity in GAIR/OpenSWE 10 days ago

Are these images publicly available?

#2 opened 10 days ago by

cppowboy

liked a dataset 13 days ago

GAIR/OpenSWE

Viewer • Updated 12 days ago • 45.3k • 1.54k • 16

upvoted a paper 13 days ago

daVinci-Env: Open SWE Environment Synthesis at Scale

Paper • 2603.13023 • Published 15 days ago • 30

liked a dataset 14 days ago

stepfun-ai/Step-3.5-Flash-SFT

Viewer • Updated 14 days ago • 1.62M • 49.7k • 285

liked a dataset 15 days ago

TIGER-Lab/WebInstruct-verified

Viewer • Updated Nov 27, 2025 • 462k • 373 • 67

upvoted 3 papers 17 days ago

In-Context Reinforcement Learning for Tool Use in Large Language Models

Paper • 2603.08068 • Published 20 days ago • 41

Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams

Paper • 2603.07392 • Published 21 days ago • 18

OpenClaw-RL: Train Any Agent Simply by Talking

Paper • 2603.10165 • Published 18 days ago • 144

New activity in Qwen/Qwen3-Coder-Next 17 days ago

Amazing , it works with open claw

#39 opened about 1 month ago by

infinityai

Yinxu Pan

AI & ML interests

Recent Activity

Organizations

cppowboy's activity

Can not reproduce evaluation results on SWE-Verified

Are these images publicly available?

Amazing , it works with open claw