SEABO: A Simple Search-Based Method for Offline Imitation Learning Paper • 2402.03807 • Published Feb 6, 2024
Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling Paper • 2502.06703 • Published 1 day ago • 76
PEARL: Zero-shot Cross-task Preference Alignment and Robust Reward Learning for Robotic Manipulation Paper • 2306.03615 • Published Jun 6, 2023
A Large Language Model-Driven Reward Design Framework via Dynamic Feedback for Reinforcement Learning Paper • 2410.14660 • Published Oct 18, 2024
RAT: Adversarial Attacks on Deep Reinforcement Agents for Targeted Behaviors Paper • 2412.10713 • Published Dec 14, 2024
BenchLMM: Benchmarking Cross-style Visual Capability of Large Multimodal Models Paper • 2312.02896 • Published Dec 5, 2023 • 1
ChessGPT: Bridging Policy Learning and Language Modeling Paper • 2306.09200 • Published Jun 15, 2023 • 9