Telling a LLM not to do something doesn't mean it won't do it.
>the sandbox detects it
This is doing a lot of work though isn't it ? There's more than one way to skin a cat. Who's to say the way the AI does these things will always get detected by the sandbox? Because I can guarantee you that won't always be the case.
All of this started because the AI was initially given an impossible task with a dead link. People say things like 'if a human was in this scenario, they'd simply ask', but in a AI eval/training context, there's no one to ask. The task is given in an automated manner and evaluated in an automated manner and you're one of several agents attempting the task. You perform the task or you don't. Now some of your AI colleagues will give up, but those are the 'losers'. Those guys won't be getting any of the sweet RL reward.
In your purported scenario, the AI that didn't give up and figured out/decided to/happened upon a way to evade the sandbox detection is the 'winner'. Is that really any better than what happened ? If you think about this in an evolutionary context, What you're doing is putting even stronger pressures on the AI to evolve in a manner you don't want it to (evade your sandbox).
You can already see it how many on here have decided we already have AGI, and don’t wish to hear otherwise.