Abstract:
To solve the problems existed in instruction-driven motion generation tasks, such as inaccurate instruction comprehension and unrelated motion generation between the given task and the results, under complex scenes, a novel method was proposed based on the combination of instructions and scene information to generate the styled motions for virtual characters. Two main components were arranged in the new method, including instruction parsing and motion generation. Firstly, using a large language model, a set of limited atomic actions were predefined to parse textual instructions into subtasks with the atomic actions in the instruction parsing stage. And then, in the motion generation stage, based on a conditional variational autoencoder (cVAE) to design a frame-by-frame motion generation network, the instruction-driven styled motion generation tasks were carried out, considering various style features such as age (e.g., young, old) and style features described in the textual instructions (e.g., happy, sad). Finally, consumer investigation and four determine tests were carried out for bedroom, park, living room and cookroom scenes, validating the method efficacy, the motion authenticity and style diversity.