Re: What's missing?

From: Robb Lovell (robblovell@intelligentstage.com)
Date: 05/02/03


The following message was posted to: dance-tech

Thanks for your input and response on this.

I am quite interested in improving my tracking software, but have been
unsure about which direction to go.  I am interested in an literature that
would be available that might allow tracking of hands or any other
distinctive features.  I agree with you that the software needs to improve,
and I only give the excuse that its that way because the domain the software
addresses has a lot of open research questions unanswered.

I think that the êyês (eyes) software that I provide at
www.squishedeyeball.com is quite good at tracking objects.  I provide many
ways to segment images, and two ways to track the objects in the resulting
images.  (that's my language for extract something different from a
background and then track it).  The software is a suite of tools for solving
image understanding problems.

Ultimately, tracking involves figuring out what in the image your interested
in following in each frame of video.  This involves four technical
challenges:

1.  Segmenting a series of images one at a time into objects that are
distinct. This is extraction of what your interested in tracking from the
image.
2.  Matching objects extracted in the segmentation step over time.
3.  Feeding in and back information in both steps to help both matching and
segmentation of future images and objects.
4.  Folding in assumptions about the environment to make contextual
decisions about what is being tracked.  This involves structuring the
environment so that the assumptions hold most of the time.  By structuring,
I mean lighting, background, costume, and movements.

While most software can do steps 1 and 2, yet steps 3 and 4 must often be
left to the user.  Its hard for a piece of software in this field to enforce
a set of assumptions about the environment, or enforce a particular way of
moving (except in gaming situations).  Often this is why the software seems
disappointing to a user unfamiliar with image understanding issues.

My software's main fault is that its too open ended.  I provide a set of
tools to solve problems in image understanding, not a set of solutions to
pre-existing environments.  Maybe I should provide a set of patches with the
software that require as a pre-requisite a particular environment and
assumptions for it to work reliably.  Even with assumptions pre-defined,
there is still the issue of initial conditions and thresholding that is left
to the user.  For instance, in tracking a hand, a system might need a learn
mode period for a particular lighting situation and distance from the camera
before it can become reliably operational.

In a way, big eye and VNS do this.  Big eye tracks color and contrast, VNS
detects movement and presence in contrasting light.  The assumptions made
here are bodies move, backgrounds don't.

Robb

A couple of notes: (sorry for the long winded nature of my answers :-)
>  There are finally just a couple of technical points that I would disagree
>  with. I don't think that the hand needs to fill 1/4 of the frame. I seem to
>  remember that the video tracking add-on to the PS2 (demoed at last year's
>  Game On exhibition at the Barbican in London) could recognise hands and
>  track them when they were much smaller than this. As I mentioned before,
>  hands are pretty distinct things, being at the end of long thin limbs (this
>  was why I chose them as an example). If the arm is just one pixel wide, our
>  system should be able to recognise the hand as the last pixel (the one
>  furthest from the body).

Perhaps we are diddling about apples and oranges on the 1/4 frame thing.
But the point is that you need some sort of resolution to be able to
distinguish a hand from a head, as well as some pretty could assumptions and
structure.

I am naturally skeptical of these kinds of demonstrations.  Having been
working in the image analysis/image understanding field for 16 years now, I
have seen my share of claims to be followed by, "well you need this kind
particular environment" and "there are these limitations" kind of
statements.  I would guess that in this situation you had to be standing in
one spot (i.e. Not locomoting around the space), you couldn't hide your
hands for too long, if you leaned over and put your had above your head you
might have your head tracked, and you had to wear clothes that were not in
the skin color range of colors, short sleeves, and the lighting was very
particular.  In a way, a system is as good as the number of violated
assumptions that it can deal with.  This system may be good at dealing with
occlusion and multiple views of the hand, but not lighting changes?  Dunno.

I suppose you could make an algorithm to track the furthest points on a
longish thing, but the assumptions made are that the arms are not covered
with cloth, and that the legs are covered.  Actually, I would be quite
impressed with a system that could track in this situation.  The frustration
is that the system only works in certain situations, and not all, bum :(.

Having said that, I am hopeful that this project can shed some new light on
this problem.  I will look for some literature on this one...

A note on my statements.  I am trying to point out that you can recognize
practically anything with the right set of assumptions about the
environment.  All of the recognition systems make assumptions in order to
function.  The problem is to make the right set of assumptions that have the
least chance of being violated in a given environment.

I have also created a system that tracks hands, feet, and head but it
required a certain set of assumptions that would work for a game environment
but not well for a dance environment.
>
>  I would also disagree with your classification of a dance performance as an
>  unstructured environment. My understanding is that in this context,
>  unstructured environment means that we can take this system anywhere, under
>  any conditions, and perform any movement, and it will track successfully. It
>  seems to me that in a performance we have substantially more control over
>  the conditions and the type of movement. This isn't to say, however, that we
>  need to set the lighting and movement on the basis of what can be tracked,
>  rather that our system should be able to be "tuned" or "taught" what we need
>  to track.

When I called the dance environment unstructured, I was making a relative
statement.  Yes to a human being, the dance environment is very structured,
but compared to normal recognition environments its fairly unstructured.  I
am mainly referring to the fact that there are bodies in the space, not to
the lighting or background elements.  Bodies, especially dancer bodies tend
to mess up assumptions that a tracking program would make to track hands.
For instance dancers regularly prefer rolling on the ground to standing.

>  I have been pondering some of the issues that you raise, particularly this
>  issue of how big the thing that we want to track. This issue has been
>  bothering me pretty much since I first tried the technology. If we leave
>  aside for the moment the issue of how "intelligent" these systems are, it
>  seems to me that maybe the problem isn't that the video tracking isn't good
>  enough, but rather that the stage is too big. Maybe it simply isn't
>  appropriate to use video tracking in a conventional performance/conventional
>  performance space. I've certainly seem more successful installations using
>  video tracking than performances using video tracking.
>
>  Also is this particular problem one with the resolution of the video camera,
>  rather than the speed or quality of the video tracking software? It is easy
>  to do the maths (size of stage/2*number of pixels) to work out how big a
>  movement you need to actually register movement reliably. In most cases this
>  gives a pretty big number which in turn affects the type of "big gesture"
>  movement that tends to be used in these performances. Again, maybe we need
>  to think about where and how we use the technology, rather than continue to
>  apply it to inappropriate environments.

You are right about this too.  Cameras are too low resolution, and computers
too slow.  If a camera could register and a computer process an image
4096x4096 instead of 320x240 at 1000 frames per second you could do a lot
more in this field.  I hand far away from the camera could still be
registered on a 100 square pixels.  The problem is that commercial systems
such as web cams, security, video, and film don't demand the same resolution
and speed as required by an image understanding problem.  Just look at the
medical imaging field, cameras are approaching 1024x1024 and the speed is a
reasonable 30 fps (sometimes).  Even so, in the medical imaging field the
recognition is of specific things in specific kinds of images.

I have often said that computers need about a 10,000 times speed up before a
lot of the image processing problems will start to be solved (processing and
acquisition speed).




----------------------------------------
The Dance-Tech mailing list has recently moved to a new address.  To post a
message, send email to dance-tech@dancetechnology.org.  To unsubscribe, send
email to lists@dancetechnology.org, with the words "unsubscribe dance-tech" in
the message body.
----------------------------------------

 



This archive was generated by hypermail 2b30 : 09/08/03