The following message was posted to: dance-tech Thanks for your input and response on this. I am quite interested in improving my tracking software, but have been unsure about which direction to go. I am interested in an literature that would be available that might allow tracking of hands or any other distinctive features. I agree with you that the software needs to improve, and I only give the excuse that its that way because the domain the software addresses has a lot of open research questions unanswered. I think that the êyês (eyes) software that I provide at www.squishedeyeball.com is quite good at tracking objects. I provide many ways to segment images, and two ways to track the objects in the resulting images. (that's my language for extract something different from a background and then track it). The software is a suite of tools for solving image understanding problems. Ultimately, tracking involves figuring out what in the image your interested in following in each frame of video. This involves four technical challenges: 1. Segmenting a series of images one at a time into objects that are distinct. This is extraction of what your interested in tracking from the image. 2. Matching objects extracted in the segmentation step over time. 3. Feeding in and back information in both steps to help both matching and segmentation of future images and objects. 4. Folding in assumptions about the environment to make contextual decisions about what is being tracked. This involves structuring the environment so that the assumptions hold most of the time. By structuring, I mean lighting, background, costume, and movements. While most software can do steps 1 and 2, yet steps 3 and 4 must often be left to the user. Its hard for a piece of software in this field to enforce a set of assumptions about the environment, or enforce a particular way of moving (except in gaming situations). Often this is why the software seems disappointing to a user unfamiliar with image understanding issues. My software's main fault is that its too open ended. I provide a set of tools to solve problems in image understanding, not a set of solutions to pre-existing environments. Maybe I should provide a set of patches with the software that require as a pre-requisite a particular environment and assumptions for it to work reliably. Even with assumptions pre-defined, there is still the issue of initial conditions and thresholding that is left to the user. For instance, in tracking a hand, a system might need a learn mode period for a particular lighting situation and distance from the camera before it can become reliably operational. In a way, big eye and VNS do this. Big eye tracks color and contrast, VNS detects movement and presence in contrasting light. The assumptions made here are bodies move, backgrounds don't. Robb A couple of notes: (sorry for the long winded nature of my answers :-) > There are finally just a couple of technical points that I would disagree > with. I don't think that the hand needs to fill 1/4 of the frame. I seem to > remember that the video tracking add-on to the PS2 (demoed at last year's > Game On exhibition at the Barbican in London) could recognise hands and > track them when they were much smaller than this. As I mentioned before, > hands are pretty distinct things, being at the end of long thin limbs (this > was why I chose them as an example). If the arm is just one pixel wide, our > system should be able to recognise the hand as the last pixel (the one > furthest from the body). Perhaps we are diddling about apples and oranges on the 1/4 frame thing. But the point is that you need some sort of resolution to be able to distinguish a hand from a head, as well as some pretty could assumptions and structure. I am naturally skeptical of these kinds of demonstrations. Having been working in the image analysis/image understanding field for 16 years now, I have seen my share of claims to be followed by, "well you need this kind particular environment" and "there are these limitations" kind of statements. I would guess that in this situation you had to be standing in one spot (i.e. Not locomoting around the space), you couldn't hide your hands for too long, if you leaned over and put your had above your head you might have your head tracked, and you had to wear clothes that were not in the skin color range of colors, short sleeves, and the lighting was very particular. In a way, a system is as good as the number of violated assumptions that it can deal with. This system may be good at dealing with occlusion and multiple views of the hand, but not lighting changes? Dunno. I suppose you could make an algorithm to track the furthest points on a longish thing, but the assumptions made are that the arms are not covered with cloth, and that the legs are covered. Actually, I would be quite impressed with a system that could track in this situation. The frustration is that the system only works in certain situations, and not all, bum :(. Having said that, I am hopeful that this project can shed some new light on this problem. I will look for some literature on this one... A note on my statements. I am trying to point out that you can recognize practically anything with the right set of assumptions about the environment. All of the recognition systems make assumptions in order to function. The problem is to make the right set of assumptions that have the least chance of being violated in a given environment. I have also created a system that tracks hands, feet, and head but it required a certain set of assumptions that would work for a game environment but not well for a dance environment. > > I would also disagree with your classification of a dance performance as an > unstructured environment. My understanding is that in this context, > unstructured environment means that we can take this system anywhere, under > any conditions, and perform any movement, and it will track successfully. It > seems to me that in a performance we have substantially more control over > the conditions and the type of movement. This isn't to say, however, that we > need to set the lighting and movement on the basis of what can be tracked, > rather that our system should be able to be "tuned" or "taught" what we need > to track. When I called the dance environment unstructured, I was making a relative statement. Yes to a human being, the dance environment is very structured, but compared to normal recognition environments its fairly unstructured. I am mainly referring to the fact that there are bodies in the space, not to the lighting or background elements. Bodies, especially dancer bodies tend to mess up assumptions that a tracking program would make to track hands. For instance dancers regularly prefer rolling on the ground to standing. > I have been pondering some of the issues that you raise, particularly this > issue of how big the thing that we want to track. This issue has been > bothering me pretty much since I first tried the technology. If we leave > aside for the moment the issue of how "intelligent" these systems are, it > seems to me that maybe the problem isn't that the video tracking isn't good > enough, but rather that the stage is too big. Maybe it simply isn't > appropriate to use video tracking in a conventional performance/conventional > performance space. I've certainly seem more successful installations using > video tracking than performances using video tracking. > > Also is this particular problem one with the resolution of the video camera, > rather than the speed or quality of the video tracking software? It is easy > to do the maths (size of stage/2*number of pixels) to work out how big a > movement you need to actually register movement reliably. In most cases this > gives a pretty big number which in turn affects the type of "big gesture" > movement that tends to be used in these performances. Again, maybe we need > to think about where and how we use the technology, rather than continue to > apply it to inappropriate environments. You are right about this too. Cameras are too low resolution, and computers too slow. If a camera could register and a computer process an image 4096x4096 instead of 320x240 at 1000 frames per second you could do a lot more in this field. I hand far away from the camera could still be registered on a 100 square pixels. The problem is that commercial systems such as web cams, security, video, and film don't demand the same resolution and speed as required by an image understanding problem. Just look at the medical imaging field, cameras are approaching 1024x1024 and the speed is a reasonable 30 fps (sometimes). Even so, in the medical imaging field the recognition is of specific things in specific kinds of images. I have often said that computers need about a 10,000 times speed up before a lot of the image processing problems will start to be solved (processing and acquisition speed). ---------------------------------------- The Dance-Tech mailing list has recently moved to a new address. To post a message, send email to dance-tech@dancetechnology.org. To unsubscribe, send email to lists@dancetechnology.org, with the words "unsubscribe dance-tech" in the message body. ----------------------------------------
This archive was generated by hypermail 2b30 : 09/08/03