Bringing the 'Minority Report' user interface to reality, without the gorilla-arm. We're a team of three research engineers developing fundamental, finger-precise hand-tracking and gesture recognition technology. We're looking for two more engineers to join us with experience in some of the following:
Computer graphics engineer:
- Solid understanding of the practical aspects of the computer graphics pipeline, shaders
- Comfortable with 3D math: vectors, matrices, rotations, projection, etc.
- Solid understanding of computer systems: caches, low-level optimization
- Game development background ideal, user interaction design a plus
- Comfortable with C/C++
Computer graphics / computer vision research engineer:
- Strong optimization or machine learning background
- Experience implementing algorithms on 3D geometry or 2D images
Very true. I work at a company building an alternative gestural input device (http://threegear.com). Here's how we have tried to address your points.
1. Gorilla arm -- keep your hands low. We support tracking and interactions literally 1cm above the keyboard / desk. We're mounting the camera above the monitor to achieve this.
2. We use gestures with built-in physical feedback. For instance, our click mechanism is a "pinch" which brings the thumb and index finger tips together. You can "feel" the physical touch event between your fingers when you trigger a command.
A couple medical device companies and a hospital are evaluating our system right now. :-)
We're actively working on supporting smaller / shorter-range sensors as well. You probably already know that in addition to the Kinect, a lot more depth cameras are on the market now: PrimeSense's Capri, SoftKinetic, PMD, Inuitive Tec, etc. All of these companies have introduced gum-stick sized sensors that can be embedded in a laptop or monitor.
Actually, it's quite possible to build a "clicker" on the Kinect. It involves mounting the sensor from above, and building completely different software that tracks the hands and fingers well.
A big problem with the LEAP is that there isn't an effective way to click / select something. Pushing forward with your index finger isn't very accurate when the finger tip is also controlling the position. Hence, you always seem to miss where you intend to click. Good selection is a pretty important piece of almost any useful application.
Disclaimer: I work at 3Gear Systems (http://threegear.com), developing technology that possibly competes with the LEAP. We solve clicking by tracking the entire hand -- not just the straight finger.
When I first arrived at MIT, I was handed a book on "How to Get Around MIT." I was impressed with the section on hacking, which included the following story about the Harvard-Yale hack:
"DKE has tried to hack the game before, most memorably in the late 1940s when they buried explosive cord in a pattern that would spell out "MIT''. Unfortunately, Harvard discovered the hack and set up a trap. They arrested several students wearing coats lined with batteries. A dean, who had been informed about the hack after the arrest, went down to bail the students out. He pointed out to the detective that the battery-lined coats were only circumstantial evidence. At this point the dean opened his own battery-lined coat and declared "all Tech men carry batteries.''"
My point is that MIT presents itself as a place that defends hacking, and it has at least been lenient in the past.
We're looking forward to working with all the new 3D camera technology coming down the pipe, and as soon as we have the resources to do so, we'll start supporting Mac and Linux too.
It's hard to specify our exact workspace, because it's the intersection of the two camera frustums. Here's an approximate estimate: 2 feet wide x 1.75 feet deep x 1.25 feet high.
What's great about the Kinect is that it lets developers go after "3D computer vision" problems rather than "2D vision."
There's a wealth of techniques from computer graphics on dealing with 3D point clouds, whereas even basic things like background subtraction are still hard (and not completely robust) in the 2D vision world.
I've just replied to the parent post with some discussion about how we differ in terms of features and technical approach.
We're also certainly cognizant of the price of our SDK. We hope that not too far down the road, we'll be able to spec out some cheaper hardware for our users.
It's indeed similar, but we think there are three important differences.
Because our cameras are top mounted, our system is more comfortable / ergonomic. You don't have to lift your hands very high to interact, and your arms are always supported comfortably by the desk. We can enable convenient interactions right above the keyboard or even turn the desk itself into a touchscreen. Basically, our system is designed to be comfortable enough to use all day.
The second difference is more speculative, because it's hard to tell exactly what the Leap can or cannot do based off their video. Our system captures the entire hand rather than just the finger tips. This lets you use more natural gestures. For example, you can spin a virtual object by rotating your hand as if you were holding it (like in the sword-waving example in our video).
Finally, we're starting with commodity hardware you can get at immediately. It may look a little clunky right now, but we're experimenting with different cameras and different frames right now. Further down the road, it'll be a monitor / laptop clip-on or even built-in. Today, we're just releasing an SDK to let developers get started.
Thanks for the feedback! All of us at 3Gear spend more time than we should reading HN.
While we're on the subject, I'm curious if you've played with matrix-toolkits-java / netlib-java. I'm considering switching to it. I've had good experiences with the Lawson-Hanson non-negative least squares that comes with netlib.
I use the Colt (Java) matrix libraries for applied math and graphics applications, but it's still quite costly compared to optimized libraries like Intel's MKL. A singular value decomposition takes about eight times as long in Colt vs MKL.
The buildings are meticulously 3-D modeled and texture mapped--you can see the texture repetition artifacts. I've read some research about automatic modeling techniques from street view images+, but I'd suspect this was done by hand.
Microsoft Research is one of the top 10 computer science research organizations in the world. Some of the work done there (such as the Kinect) is game changing and will continue to be a resource Microsoft can draw upon.
As an MIT PhD student, I know many people who are eager to work at Microsoft research.
It is easier to acquire specialized skills in computer vision and machine learning in grad school than in industry. I've also personally had a good experience incubating technology during my PhD, although this depends on your advisor and your country. The PhD programs in the US are usually longer than the EU (6 versus 3 or 4 years), but typically provide more freedom and less pressure to publish constantly.
The levamisole test kits should be made available to the smugglers. If the smugglers knew they were shipping (detectably) impure cocaine, there would be pressure on the producers to stop cutting the product.
(... and a little voice in my head tells me that we'd be better off if drugs were regulated rather than illegal.)
The author, Aaron Swartz, harps on the fact that the film doesn't show enough teaching and instead "hides behind charts and graphs." The author would rather the film show "terrified kids up on the big screen." This seems to me like favoring anecdotes over data.
The author also attributes the crisis in American education to standardized tests. He doesn't back this up. Standardized tests work well in countries like China and India, from which so many of our engineering grad students hail. Granted that no standardized test is perfect, it boggles my mind that Swartz would call it the educational crisis.
Bringing the 'Minority Report' user interface to reality, without the gorilla-arm. We're a team of three research engineers developing fundamental, finger-precise hand-tracking and gesture recognition technology. We're looking for two more engineers to join us with experience in some of the following:
Computer graphics engineer:
- Solid understanding of the practical aspects of the computer graphics pipeline, shaders
- Comfortable with 3D math: vectors, matrices, rotations, projection, etc.
- Solid understanding of computer systems: caches, low-level optimization
- Game development background ideal, user interaction design a plus
- Comfortable with C/C++
Computer graphics / computer vision research engineer:
- Strong optimization or machine learning background
- Experience implementing algorithms on 3D geometry or 2D images
- Solid understanding of computer systems
- Research experience ideal
- Comfortable with C/C++
Contact: [email protected]