This paper addresses the difficulty of assigning credit across long, multi-step agent trajectories during reinforcement learning, where sparse end-of-episode rewards make it hard to identify which actions actually mattered. DRACO introduces dynamic, fine-grained rubrics that score individual steps within long-horizon agent rollouts, providing denser and more accurate training signal than trajectory-level rewards alone. The approach improves training stability and downstream task performance for long-horizon agents compared to standard sparse-reward reinforcement learning baselines.