The first news item in point of time relates to my thesis work.
In brief, I'm trying to demonstrate how to track the human torso in monocular (single-camera) video sequences. Take a look at some latersections for more technical (read "dorky") details.
For the most part, I feel like I've executed my thesis work to expectations that have not been very well defined for people who seem to have little to no interest in the subject. I defended in May to a less-than-impressed audience, was passed because my advisor doesn't really advise me at all, and was required to demonstrate some new results to complete the procedure.
Now it's down to waiting for them to finish reviewing the thesis, a process that has required me to make two trips to Columbus for roughly three hours of meeting time (in total). So that's 28 hours of driving (seven each way times two trips) for about three hours of grammar edits and results discussion.
The deadline for all of this to be done is September 20th, and I'll officially have my degree in December. If they drop the ball on this, though, I probably won't finish until I can get someone else to cough up the dough. OSU paid for my current stint, and I can't afford it on my own.
Paying my own bills is best left for another entry--so here's the details I promised above.
Tracking
Tracking, as a rule of thumb, is a difficult problem. Generally, it's a search problem: given an image, search it to find an object of interest. For example, given a picture, you want to find the human torso in it. The brain, of course, does this routinely and pretty easily, but n'er-do-well computers have a difficult time wrapping their heads around it.
One reason is because computers don't know what people look like or how they move. And unlike blocks or ball bearings, humans move in complicatedways.
The second reason is because computers cannot integrate many parts of the picture together. The brain perceives the image as a whole, not a discrete grid of pixels, so we have an easier time putting pieces of thepicture together.
Current Solutions
Most current solutions to the problem require some form of probabilistic formulation: it looks like we found the object here in the prior frame in the sequence, so it's probably around that same area in the next.
Using this kind of solution is usually fairly efficient but has a couple drawbacks. First, the computer must learn two things ahead of time: what the object looks like and how the object moves. This means that the tracker must maintain a pretty specific model for what's being tracked,and that it must know ahead of time what it will be tracking.
While the tracker can usually track very well given that information, it isn't very nice when you want to track novel information right away.(Learning processes take time and incur startup costs.)
My Idea
Using some software developed by a group at Ohio State, I have formulated a solution that analyzes the contour of the figure to be tracked and breaks it into segments. A human, for example, might have six segments: a head, two arms, two legs, and a torso. (Flashback to Pinky and theBrain here: oooo, Brain, a monster! It has ooooo, two arms!)
By using characteristics like area and location for each segment, I think we can identify segments in each frame with a high degree of probability. In other words, I can say, "Segment 2 in Frame 96 is the same as Segment4 in Frame 97."
Coupling this with a model would provide some semantic detail, like, "Segment 2 has the largest area, so it's probably the torso." As soon as all of the pieces fit together, you have a tracker that does not require any real ahead-of-time learning. Instead, it uses what it's tracking to provide results, which is closer to how the brain functions (I think).
Alright, it's time to complete the dorking. I hope you enjoyed it. Later I'll talk about what I've been up to more recently--the job hunt.
As always, devotedly yours,
E
No comments:
Post a Comment