Reward only on the final answer and learning crawls; shape intermediate tool calls and the model learns to spam SQL. Use outcome-grounded credit assignment, GRPO over a group of trajectories, and fine-tune a pre-aligned model so you're not teaching tool syntax from scratch.
20 answers per topic instead of 10, plus saved progress and bookmarks · no cardor unlock all 523 remaining answers · ₹2,000 / $25
