hyeong
Posted on Jul 16
• Originally published at dbhyeong.github.io
#llm
#agents
Anthropic Frontier Red Team이 언어 모델 12종에게 진짜 로봇을 쥐여 주고 측정한 Embody 벤치마크를 뜯어봤다. 관절을 직접 몰게 하면 대부분 실패하고, 사전 학습된 제어기를 감독하게 하면 갑자기 유능해진다. 결론은 '능력은 모델이 아니라 접근
hyeong
Posted on Jul 16
• Originally published at dbhyeong.github.io
#llm
#agents

Imagine working at a warehouse or office sometime in the near future, and you’re asked to help a new trainee learn the basics of…

Data standards and real-world deployment remain key barriers to progress.

decided to share my recent observations from working with AI coding agents / LLMs everyday...

Introduction Every team building with AI agents hits the same wall. The demo works...

When Tianyu Li, a doctoral student in Mechanical Engineering and Applied Mechanics (MEAM), and George Jiayuan Gao, former…

I've been running a small agentic eval harness against a local model and I'd like a sanity check on...