Every modern assistant demos well, because language models are uniformly good at sounding right. The demo tells you almost nothing about behaviour on your content and your edge cases.
What differs between products is grounding, boundaries and handover — none of which is visible in a scripted demonstration.
Where do answers come from — your documents and records, or the model's general knowledge? Can it cite sources? And what does it do when retrieval finds nothing?
An assistant that answers from general knowledge about your product will be confidently wrong in a way that is very hard to detect until a customer complains.
The failure paths. Ask it things it cannot know and watch what it does.
For anything beyond policy questions, yes — and then permission resolution becomes critical.
Judging on the demo. Every product demos well; they differ on grounding and handover.
Half an hour with your own data usually saves reading three of these. The guides will still be here afterwards.