Voice as sole input mechanism for HCI is too limiting in a number of dimensions: determinism in what you can express, efficiency as well as discoverability (e.g on-screen affordances), and doesn’t work in a large number of environments (generates noise).
Moreover, since the results or output is typically displayed on a screen, you’d be missing the opportunity if you don’t take advantage of on-screen affordances.