Unless it's CUDA only, a fair amount of software now supports AMD ROCm. And if you're already wanting to write an inference engine, 32G will be more than fine.
I've had a R9700 and it'd run the lower quants (4-5 bit?) of Qwen 3.5 27B reasonably attached to Hermes Agent. Not quite a 3090, but still respectable speeds.
Of the two affordable options for good 32G compute, the R9700 would be the best of the two (compared to the Arc Pro B70).