Sushil’s Karma: Notes on AI Persona Design

There is this company I did an AI/UX audit for, which provides a virtual AI secretary that takes all your calls. If it is relevant, it passes the call onward to your phone. It detects spam and unwanted calls and saves you from all that hassle.

When I called the company’s phone number, the sound that I heard was very human. Indian female, North Indian accent. She told me the features and the pricing plans, with ambient noise around her and keyboard sounds. I have written another post about this particular aspect — the typing sounds and the waiting thereby. However, I realized this: I, as a user of this product, want an AI agent. I might not want Sunita from Delhi. Of course, if someone called my number, they should hear Sunita. Many of those callers do not know that such AI callers exist. They should believe I have a human secretary working in an office, with a computer in front of her.

What if Sunita had told me this: “I will be the one to handle any calls to you. You could also choose Kavita or Sunil. Would you like me to switch the call to them?”

But there is another scenario. What if I do not need a human persona as a secretary? I am subscribing to the service of a virtual AI secretary, is it safe to assume I will always want a human voice? What if I want a non-human, robotic voice, like R2D2? What would be the benefits of this, one could ask. Assistants have always been at the receiving end of their boss’s emotional outlet. If a client makes them angry, that suppressed anger comes out on assistant Sushil. Though ideally this should not be the case, what if someone wants a virtual AI secretary just because they don’t want to get more bad karma by abusing their human help? Scolding a robotic voice may be preferable to some. This is just one scenario which came to mind, absurd as it may well be. The question is whether anyone has done any studies to find out more. Why is it always different types of human sounds being shown in voice personas?

Of course, almost all AI voice agents have plenty of different voices in different accents. But should this be the first real choice, before deciding between those. The very first choice should be this — do you need a human voice at all? Many other things are connected to this. If I know my secretary is virtual, is the fake typing sound really necessary? If the latency is due to technical constraint, then it makes some sense at least. Instead of saying “please wait, fetching the info”, they fill in the silence with typing. If the latency is part of the agent persona, simulating a human finding data, then it has no place if I know my assistant is AI. Since an AI assistant would know most information, at least relevant to me, almost instantaneously. Why would I spend money to hear my AI secretary typing without writing anything?

There should be ambient noise though, to make sure the call hasn’t dropped when there is a moment of silence in the conversation due to technical latency. Maybe I should be able to choose between white noise, green noise, or black noise, because I may be hearing that noise for a substantial time period every day. That sounds too impersonal, though. This particular case has a voice agent who is your personal assistant. Would it be ok, if he asks, other personal questions like a real assistant might? Just like in the real world, when a Boss and an assistant are both waiting for something else, and the assistant goes, “How was your daughter’s violin recital yesterday?”. Though it may sound irrelevant, it gives the assistant to know the person better, helping him, help him better. But that is a much longer and separate discussion, and suitable for a different post altogether.

{If you would like to get a UX audit of your voice AI service, feel free to contact anand[at]rega.in}

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *