Virtual Human Animation
Jorge Gonzalez Hardman February 17, 2011
Abstract Animation has been with us for over a century, changing and evolving from traditional 2D animated sequences to 3D environments. Throughout this history, animations have often represented human beings. Humanoid characters have been used throughout the ages to amaze, entertain, and even teach. 3D environments have allowed the creation of ever more realistic human characters, using a variety of technologies available. One of the biggest problems is the amount of detail, movements, and mannerisms that humans have, and while very human-like avatars have been available for a while, there is still a long way to go to have a human character that is completely indistinguishable from a real person when thinking over the purely visual.
Introduction.
As Badler [Bad97a] mentions, and still holds true today, human animation is ever growing in speed and quality, getting better results with more powerful technology that is available in todays world. Virtual humans are something that may be used in various tasks, substituting real humans to be able to perform these task before putting forth the eort of building a working prototype, in other words seeing how humans will act under certain conditions. An example of this would be the use of virtual humans to plan a building evacuation route in case of re or other such disasters. Other uses for virtual humans are those of creating avatars of living humans, useful in several areas, most notably in the entertainment industry, where they may be used to represent characters in situations where using real humans may be too dangerous. Another huge area in which the use of human avatars is of enormous contribution is in medicine. Representing humans in such an area requires very accurate representation of the human body, and may be used in simulation and training. Such are the uses of animating men and women, and to this end there are several techniques and standards used. This paper will take into consideration a brief overview of the history of animation and virtual humans, before focusing on these techniques, in particular inverse-kinematics and retargetting animation. In brief, retargetting animation is an extremely useful way of translating motion from one character to other characters that have dierent features, thus allowing the animator to save time. This makes use of inverse kinematics to solve key poses and give a uid motion on dierent characters using only one generalized animation. The dierences between the characters may be enormous, and may still translate the motion to t the new characters body. An excellent example of retargetting animation occurs in games such as Spore, which 1
allows the user to create their own characters in almost any shape imaginable, and makes use of this technique to animate these characters, allowing them movement in a variety of ways, no matter what bone structure they may have, as described by Hecker, Raabe, Enslow, DeWeese, Maynard, and van Prooijen [HRE+ 08]. As mentioned, this makes use of Inverse Kinematics, which plays a key role in 3D animation, acting as both a tool to make movement more realistic, as well as easing the process for 3D artists to create their characters and position them as needed. For example, an artist may position extremities without the need to manipulate every joint in the characters limbs, by moving the hand alone, and having an algorithm automatically generate the rest of the positioning, angles, and movements.
2
2.1
A Brief History
Origins
As Webber, Phillips, and Badler [WPB99] explain, the use of virtual humans came about in several dierent elds through independent means. The one that jumps to mind, of course, is in the eld of entertainment, as even before three-dimensional characters were available, two dimensional animation had been long standing in this eld. Edward George Lutz wrote his book titled Animated Cartoons; How They Are Made, Their Origin and Development. This would become the rst textbook in what would be modern traditional animation. The rst example of what the general public would concede as two dimensional animation is Disneys famous Steamboat Willie in 1928, which managed to synchronize the visual animation to the audio track, completing the eect. Since then animation has gone through several changes in techniques and methods to bring about the nal product, but the basic principles of animation have stayed relatively untouched. Early 3D modelling and animation techniques were developed at the University of Utah, Ohio State University, and the New York Institute of Technology. Webber, Phillips, and Badler [WPB99]. Other uses for virtual humans arose at around the same time in other elds. Virtual humans were developed for use in crash testing, athlete performance analysis and training, motion analysis and notation, as well as for understanding motion. Noting these uses of these rst developments in virtual human animation, the major hurdle that was being overcome at the time was that of human motion. Studying, identifying and replicating this motion brought about these rst examples of virtual humans and understanding these motions and using them to gain data. Since, more and more data has been added to the understanding and representation of humans in the virtual environment.
2.2
Virtual Humans Today
Today there are several techniques, standards and technologies that aid in the representation of virtual humans. Motion capture helps capture movements as made 2
by actors that translate into realistic motion representation. Animation is being separated from individual characters and may be translated for use on other virtual avatars, and the representation of these virtual entities is concentrated on not only the replication and understanding of motion, but on the replication of human behaviour outside of the physical world, such as their mannerisms. A good example for virtual human animation today is that of Microsofts Milo, demoed during the 2009 Electronic Entertainment Expo. Microsoft Milo represents a young boy who simulates a real boy, in a virtual environment, who can interact with a player through video and audio, understand in and reacting to speech. The video demo comments on the use of several techniques, among them recognizing movement and voice, to have the virtual character react to the human players actions. The virtual character does not only go through the motions of a normal human boy, but also represents human mannerisms, down to emulating human emotions, such as embarrassment.
3
3.1
Modern Techniques
2D Techniques
2D animation has mostly kept its basic principles throughout the last century, while applying new methods of building the animations. Computer assisted animation brought about the use of automatic in-betweening. This allows animators to create key frame drawings, and exploiting the animation software to automatically create the frames in between these keys by rotating, translating and morphing the object to create a smooth animation. 2D animation however is done while keeping a 3D world in mind so as to represent depth of objects in the scene. Keeping in mind that while 2D animation represents only two dimensions, the animator works in a 3D space, as characters and objects in the 2D scene will be rotating along the x, y, and z axis. As Di Fiore, Schaeken, Elens, and Van Reeth [DFSEVR01] point out, rotating along the z-axis leaves no problem, as the object in the 2D plane is simply rotated, and does not change its shape as viewed from a two-dimensional perspective. Rotating on the x and y axis however changes the overall shape of the object being drawn. By working with vector drawings and layering the animator may create 2D drawings within a 3D space, and may use forward and inverse kinematics discussed in further in section 5.2 to help with movement and positioning of scene elements.
3.2
3D Techniques
3D animation on the other hand is modelled and shaped in a 3D environment using 3D objects. Moving characters and objects are often represented as a set of articulations, where a set of pieces or segments form the object. Through the use of inverse kinematics (further discussed in section 5.2) movement from one segment will aect connected segments, and the according transformations and morphing of connected segments are calculated. As Watt, A.H. and Watt, M. [WW92] discuss, soft object animation allows deformation to occur, which is of tremendous use when applying 3
the basic principles of traditional animation, such as squashing and stretching, which would otherwise be impossible for rigid three dimensional objects. Motion for said objects in animation packages present animators with similar layouts expected in 2D animation. Key framing may be used to create movement for the object or character, with the software creating the in between frames. For smoother, more natural movement however, motion capture technology may be used to create motion, but this has the disadvantages of not capturing or adding any exaggerations that stem from the principles of traditional animation, and must be added in later if desired. Movement may also be specied by splines, which also involves key-framing, and involves specifying a path in which the object will move while also keeping velocity in mind. Several techniques are used in animation, among them is retargetting animation, specied in section 5.1, and many of these are shared by both 2D and 3D animation environments.
4
4.1
Standards
H-Anim
The human body consists of anatomical segments (such as upper arm, forearm, hand, skull, hind foot...) which are connected to each other by joints (shoulder, elbow, wrist, skull base, ankle...). Preda, M. and Preteux, F [PP01].
These segments and joints then relate to each other by dening a root node in the character (which in humans would be represented by the pelvis) and building o of it. The root node has child nodes, which are aected by their parents movements, and in turn have their own child nodes. Thus, taking a human character as example, when moving a forearm, the hand will move with it, but legs will not be aected. Segments are what forms the character, while joint segments provide the ability to rotate and move these segments, which deform the segments as needed to represent a natural movement. However, if a character mesh is not correctly attached to the segment, it may result in extreme deformations that do not faithfully represent the objects movement. A huge advantage that comes with this representation however, means that any object that works with the same segments and articulations can then re-use animations, thus reducing the amount of work when animating several characters.
4.2
MPEG-4 SNHC
MPEG-4 SNHC (Synthetic and Natural Hybrid Coding) is a part of the MPEG-4 standard that began to form in 1993, and was dened as a standard by 1999. SNHC represents a scene as a collection of audiovisual objects which can occupy points in space, and may move with time. This standard allows the integration of both synthetic and natural audio or visual data. When relating this to the human body, it comes back to dening the human body as a set of parts or segments that are connected by articulations. MPEG-4 SNHC, as pointed out by Thalmann, D. and Vexo, F [TV03], denes facial and body animation to describe an object and animate it. Within this denition, four sets of parameters are dened, which include the denition of facial and body geometry, as well as the denition of facial and body animation. Body 4
denition parameters are thus based on the specications of H-anim, allowing the creation of a character through the use of segments and joints in the same way, and allows the same hierarchical relationship for joints, which are joined by the segments which dene the geometry between the joints. Segments can then be used along with inverse kinematics to help animate the character. Animation parameters then are the denition of movements made by these joints and segments, by rotating or translating the joints, which subsequently aects the position and angle of child joints, along with the modifying the mesh to match the rotations. Facial denition parameters are similarly made to model facial characteristics by dening facial points, which may then be animated through the use of facial animation parameters.
Techniques for Character-Independent Animation
One of the largest advantages of computer animation, as opposed to traditional means, does not only revolve around the use of automatically generated frames, and motion capture, but the ability to reuse this information and apply it to other characters and models, keeping the general motions in place, while applying them to dierent models, that may vary in size or shape, but contain similar structures. This allows animator to save huge amounts of time and money. By building 3d objects as structures, as seen in both H-Anim and MPEG-4 SNHC, and through the use of either inverse or forward kinematics animation of a character is made easier, while retargetting animation allows to translate these motions to other structures.
5.1
Retargetting Animation
As described by MPEG-4 SNHC and H-Anim, body structures are dened, and movements are then applied to them. These standards then keep the structure of a 3D object separate from its animated movements, making these actions independent from the geometrical structure and applicable to other similar structures. Retargetting animation then saves both time and money, as stated before. This motion retargetting then is applicable to structures with the same number of joints, but where the segments may vary in dimensions. This means then, that motion must consider these dierent dimensions when applying to the new structure. As pointed out by Gleicher [Gle98], a character that is walking or running along must be touching the oor, and a dierence in limb length would change the length between the pelvis and the ground. Thus, the height of the pelvis is stored relative to the position of the feet. Gleicher [Gle98] denes this as a constraint, one of many which when translated from one structure to another must be adapted so as to not be broken. This way movement from one character may be correctly adapted to another. Inverse kinematics may be used to nd adapt the movement of each frame in the animation. This then is also very useful in both animations for bodies and for faces. To give an example from the gaming industry, a very identiable application of retargetting animation is used in the game Spore. Users may create their character in game, based on a spine, on which every vertebrae acts as a deformable piece of the
body, to which several dierent limbs may be attached then deformed, by stretching to make them longer, or adding muscle mass, thus giving the player the ability to generate a very large amount of distinct creatures, which may then walk around, jump, y, attack, sing, among other animations that are possible. This means that the animation must be applied to a geometry that is not completely known before hand. Several techniques can be used, but the basis is that limbs and body parts are treated as separate bodies. These bodies are then assigned by calculating positions and rotation of the bodies. The position and rotation of a given body bi on a specic character in character-relative Euclidean space (i.e. what is displayed in the Spasm view-port) is referred to as the specialized coordinates or pose of the body, qsi. The function G takes the specialized pose to the body-independent generalized position and rotation coordinates, qg, by qg = G(bi, qsi,m). Similarly, given qg and a body bi, the function S produces the qsi for that body, qsi = S(bi, qg,m). Hecker, Raabe, Enslow, DeWeese, Maynard, and van Prooijen [HRE+ 08]. They then go on to use G to nd out how body parts that are attached to the moving part will react. These calculations along with several other techniques applied to the movements, such as nding how movement varies depending on the position and rotation of the body, and mirroring, and culminate in retargetting animation at runtime for distinct structures.
5.2
Inverse Kinematics
Finally, Inverse Kinematics, that are widely used in several of the methods and techniques mentioned before, is the process of nding out the joint angles on a structure, given what is called an end eector. The position of this end-eector is dened by the user (an example would be an animator positioning his characters hand at a higher location), after which calculations are made to nd the joint angles and thus position the structure in such a way that it may reach the desired point naturally. Badler [Bad97b] mentions that inverse kinematics does not describe human perfectly, but works as an aid to nd the right angles for the articulated gure. Other problems may arise from using inverse kinematics. As structures grow more complicated, simply using an IK algorithm will become slower to process, and be in more danger of failing to nd a solution, and may also generate problems when not taking any measures to avoid collisions between segments, or fail to nd a solution when to many of this type of constraints are applied. as pointed out by Tolani, D. and Goswami, A. and Badler, N.I. [TGB00]. This being said, by including the right constraints, inverse kinematics plays a large role in positioning articulated structures correctly. Taking into example a limb, say an arm, the hip, knee, and ankle will act as joints. Considering this example in two-dimensions, so as to simplify, each of these joints will then present movement vectors and angular velocities, provided by rotation along the z-axis. To nd an inverse kinematics solution, the Jacobian is used. The Jacobian J is the n x m matrix of partial derivatives relating dierential changes in X, written as dX. Watt, A.H. and Watt, M. [WW92]. This becomes useful when nding a solution for inverse kinematics, as it involves nding the inverse of the Jacobian. To nd the Jacobian for use in the calculation of an IK solution, it is possible to calculate the velocity at the end eector, by obtaining the sum of the cross products of each of the structures 6
angular velocities by their respective position vector travelling from the joint to the end eector. By calculating this it is then possible to nd the Jacobian, which may then be inverted to nd a solution. Having done this however, it is very possible that no solution may be found, in which case it is determined that the end eector simply can not reach the specied coordinates with its structured articulations. On the other hand, and especially if the calculations are made with too great a distance between one another, it is likely that more than one solution is found. An example of this is reaching for a point directly in front of you with your hand. If you do not require to stretch your hand to the limit, then there is a very large amount of solutions in which your arm can bend to reach the point (elbows out, relaxed at the sides, among others). To limit this, calculating the positions at smaller intervals makes it easier to nd a unique solution. Watt, A.H. and Watt, M. [WW92] state that the growing complexity as more and more joints are added, along with the freedom that other methods based on forward kinematics provide, make inverse kinematics less attractive to animators compared to forward kinematics, though computer animation tools often allow for both forward and inverse kinematics to allow the animator higher degree of freedom when animating, letting them simply move characters limbs or end eectors and positioning the characters limbs accordingly.
End Notes
The creation of virtual humans has many applications, and throughout this paper several of the methods, techniques, and uses for them has been explored. Emulating human behaviour is of extreme use not only in the entertainment industry, but in the scientic and engineering communities too. As computers become more powerful, more accurate representations of human beings can be made, which makes creating a truly authentic virtual human a challenge, even in this day and age. Not only do we have them emulate our movements, but also get them to emulate our behaviour, which adds to the challenge.
References
[Bad97a] [Bad97b] [BCL02] [Chr05] N. Badler. Virtual humans for animation, ergonomics, and simulation. nam, page 0028, 1997. N.I. Badler. Real-time virtual humans. In The Fifth Pacic Conference on Computer Graphics and Applications, 1997. Proceedings., pages 413, 1997. S. Battista, F. Casalino, and C. Lande. MPEG-4: a multimedia standard for the third millennium. 1. Multimedia, IEEE, 6(4):7483, 2002. Chris Webster. Animation The Mechanics of Motion. 2005.
[DFSEVR01] F. Di Fiore, P. Schaeken, K. Elens, and F. Van Reeth. Automatic in-betweening in computer assisted animation by exploiting 2.5 D modelling techniques. In Proceedings of Computer Animation 2001, pages 192200. Citeseer, 2001. [Gle98] M. Gleicher. Retargetting motion to new characters. In Proceedings of the 25th annual conference on Computer graphics and interactive techniques, page 42. ACM, 1998. M. Gleicher. Comparing constraint-based motion editing methods. Graphical models, 63(2):107134, 2001. K. Grochow, S.L. Martin, A. Hertzmann, and Z. Popovi. Style-based inverse kinec matics. ACM Transactions on Graphics (TOG), 23(3):522531, 2004. C. Hecker, B. Raabe, R.W. Enslow, J. DeWeese, J. Maynard, and K. van Prooijen. Real-time motion retargeting to highly varied user-created morphologies. ACM Transactions on Graphics (TOG), 27(3):111, 2008. H. Ko and N.I. Badler. Animating human locomotion with inverse dynamics. IEEE Computer Graphics and Applications, pages 5059, 1996. J.S. Monzani, P. Baerlocher, R. Boulic, and D. Thalmann. Using an intermediate skeleton and inverse kinematics for motion retargeting. In Computer Graphics Forum, volume 19, pages 1119. Wiley Online Library, 2000. F.C. Pereira and T. Ebrahimi. The MPEG-4 book. Prentice Hall PTR Upper Saddle River, NJ, USA, 2002. M. Preda and F. Preteux. Advanced virtual humanoid animation framework based on the MPEG-4 SNHC standard. In Euroimage ICAV 3D 2001 Conference, 2001. J. ROSSMANN and C. SCHLETTE. The Simulation and Animation of Virtual Humans to Better Under-stand Ergonomic Conditions at Manual Workplaces. R.W. Sumner, M. Zwicker, C. Gotsman, and J. Popovi. Mesh-based inverse kinec matics. ACM Transactions on Graphics (TOG), 24(3):488495, 2005. D. Tolani, A. Goswami, and N.I. Badler. Real-time inverse kinematics techniques for anthropomorphic limbs. Graphical models, 62(5):353388, 2000. 8
[Gle01] [GMHP04] [HRE+ 08]
[KB96] [MBBT00]
[PE02] [PP01] [RS] [SZGP05] [TGB00]
[TV03] [WPB99] [WW92]
D. Thalmann and F. Vexo. MPEG-4 Character Animation. KI, 17(4):39, 2003. B.L. Webber, C.B. Phillips, and N.I. Badler. Simulating Humans: Computer Graphics, Animation, and Control. 1999. A.H. Watt and M. Watt. Advanced animation and rendering techniques, volume 104. Citeseer, 1992.
[]