You still need to do the leg work. Even when using LLMs to create deterministic tools you need to verify that these tools indeed perform the correct actions, and not just do something which statistically is often right. Using LLMs to create deterministic tools is better than just prompting LLMs with markdown files hoping the result will be the same (still legally and ethically problematic). However, if that deterministic tool is blindly maintained with LLMs the new releases have the risk over losing its deterministic behavior over time.
As Snyk showed In their VulnBench report, auditing with LLMs is not really reliable: