Monday, February 14, 2011

GEA

I'm sorry that I have not updated this blog for so long, there are many ideas to try out! But I have decided that at least I must include the screenshots of the prototypes I am working on, if only to remember them in the future.

We are happy to have a paper accepted at CVPR 2011, so I will start with the screenshot of an old friend:
This is obtained by a Haskell implementation of the GEA (global epipolar adjustment) method proposed in the paper. It is faster than sparse bundle adjustment, and gets a solution which is very close to the optimal one. Here are two more 3D reconstruction examples from point correspondences taken from bundle adjustment in the large:

Trafalgar



Dubrovnik

Thursday, May 28, 2009

SiftGPU

The library now includes an interface to the excellent ChangChang Wu's GPU implementation of SIFT. Vision problems that were recently considered infeasible can now be easily solved in standard computers.

With a NVIDIA 9800 GT we obtain interest points and descriptors at ~ 25 fps for 640x480 images, depending on the level of detail in the scene. Fast GPU point matching is also provided.

(Using a 8600M GS we get ~10 fps. Most of the time is spent in the smallest scales, so we can get a big speedup if we start from octave 1.)

To appreciate the computing power of SiftGPU you should run the detector on a live camera. Here are some screenshots of a simple object recognition prototype. First, of course, the H&Z book:

Then Real World Haskell, detected just from a leg of the insect:



This works very well, but the number of known objects is small. We need a fast method for feature indexing.

SiftGPU requires specific hardware, but it is extremely useful. Clearly, easyVision must be finally cabalized so that we can install a basic, portable system, and any desired optional modules.

Monday, January 12, 2009

Tutorial

After a long delay, I have finally released a first version of a tutorial for easyVision. I know that this system is only useful to me, but trying to explain something to other people is the best way to find things that can be improved, have a poor design, or more probably, that I don't really understand.

For instance, I noticed that the definition of camera combinators was too verbose. We can now write more elegant programs like this:

main = run $ camera >>= observe ~~> f . g . h >>= zoom

In Haskell, if a program compiles it is very likely correct. Also, ugly code strongly suggests that the essential problem is not well understood. Trying to write a better solution we often discover that it can be reformulated in a more general way. And this opens up still more possibilities for improvement.

Saturday, June 28, 2008

Interest Points

Recent algorithms for detection of scale and rotation invariant interest points have produced fantastic advances in many computer vision fields. For instance, explicit segmentation is no longer required in some object recognition applications.

Interest points are typically computed as local extrema in the scale-space of an appropriate response function. The celebrated SIFT by D. Lowe is based on the Laplacian, which is approximated by differences of gaussians in an image pyramid.

I'd like to include in the easyVision library a reference implementation of a good interest point detector. Some experiments indicate that the determinant of the Hessian obtains very satisfactory results. In a first stage I am not very worried about speed, so we will explicitly work with the full image smoothed at different levels, without any downsampling.

The following screenshot shows the interest points detected on a 640x480 image by analyzing 15 scales up to sigma=25. Blobs of different sizes are detected with excellent accuracy.

Woody Allen's Interest Points

This extremely unoptimized algorithm is however quite fast. With half-size images (320x240) we get more than 15fps in a modern laptop, and at 640x480 it runs at 3fps. This is of course not real-time, although it could be acceptable for some live video applications. In any case, note that the program currently runs single threaded. Most of the time is required by the Gaussian filters, which can be done in parallel. (This is the next step I will try to solve, but first I must consolidate a few things in the library.)

A simple local descriptor based on local gradient distributions is also computed for each point, which can be used for object recognition and image matching:



(The current implementation of the rotation invariant descriptor, based on explicit image warping, is quite slow. In the above examples we work with a scale-only invariant descriptor.)


I am currently preparing more detailed installation instructions and other information about the library.

In a future post I will describe some low level algorithms implemented directly in Haskell.

Saturday, January 12, 2008

3D point sensor

The position of a point in space can be computed from two views using triangulation:


Of course, to get a reasonably accurate 3D reconstruction (modulo scale) we need the parameters of the stereo pair (focal lengths and relative position and orientation of the cameras). Otherwise the estimated 3D points may suffer unacceptable projective deformation.

One of the most wonderful results of multiview geometry is that the calibration parameters can be inferred automatically just from the views of a small number of arbitrary 3D points (in unknown positions).

The easyVision example program "demostereo" illustrates this idea using two webcams. The interest point can be detected by a simple heuristic on the hsv color space. And a good estimation of position and velocity is obtained by a Kalman filter. First we adjust the parameters of the region detector and check that the tracker works:

If the point "stops", it is stored. When the set of points produces a promising estimation of the stereo geometry, the camera parameters are updated in the 3D view:

Subsequent estimations should be progressively closer to a similar reconstruction.

Now we can register points in the desired positions. For example, here we mark something like a box:


Note that auto-calibration from a single fundamental matrix is possible here because the camera model has only one internal unknown parameter, the focal distance.

This program shows in a very attractive way the basis of stereo vision. And this kind of 3D point capture method can also be useful in some applications.

Friday, December 21, 2007

simple object recognition

Under controlled conditions, simple objects can be recognized using just elementary texture and color features. For example, we can recognize playing cards from histograms of local binary patterns and hue values.

First we must locate the cards in the scene. To do this we can use any method which is able to find rectangles with the right aspect ratio:

Then we rectify the perspective views:

And finally the feature vector which characterizes each card is computed from the rectified view. Classification is based on the nearest neighbor method:


(This is just an example used to test some camera combinators in the library. Cards can be better recognized by many other methods.)

In fact, LBP histograms can be directly used to characterize more complex, non planar objects. There is a simple illustrative easyVision program (simpleclassifier.hs) for recognition of objects in a small catalog. New prototypes can be added with a mouse click...

Monday, November 19, 2007

camera combinators

Lazy evaluation and higher order functions are the basis of a very useful programming abstraction. Using standard Haskell list functions we can easily define combinators which work on the infinite sequence of images captured by a camera.

A camera is just an action which returns an image:
Type Camera = IO Image
A camera combinator is a function which typically takes one or more cameras and other arguments and produces a new camera. For instance, we can fuse the images obtained by two cameras into a "panoramic" view:
panoramic :: Camera -> Camera -> IO Camera
Then any program using this camera automagically receives frames like this:


Pedro and Antonio working hard in our lab

(Currently the required synthetic rotations must be adjusted manually, but we are working on an automatic method. And, of course, the cameras must be very close to each other for this to work.)

We can now create a panoramic view by joining an arbitrary number of cameras by something like this:
pano <- foldM panoramic (head cams) (tail cams) 
(I will try to prepare a demo with 3 or 4 cameras...)

The real power of camera combinators in Haskell is the fact that we can work with the infinite lazy sequence of the images captured by the camera. We have defined a "virtualCamera" function which adapts ordinary list functions to work with the IO list of images. Using it we can define, for example, a camera which "intercalates" average frames:
interpolate = virtualCamera (return . inter)
where inter (a:b:rest) = a : x : inter (b:rest)
where x = 0.5 .* a |+| 0.5 .* b
We can also compute a weighted average of the sequence:
drift alpha = virtualCamera (return . drifter)
where drifter (a:b:rest) = a : drifter (x:rest)
where x = alpha .* a |+| (1-alpha) .* b

For instance, the following effect:


can be obtained by the following composition of combinators:
cam <- getCam 0 size
>>= monitorizeIn "original" (Size 150 200) id
>>= asFloat
>>= drift alpha
>>= interpolate
The "worker" function receives just a simple IO ImageFloat:
worker cam win = do
inWindow win $ do
cam >>= drawImage
But each call automatically performs all the computations defined above.

We have defined camera combinators to filter frames very different from the previous ones (movement detectors), detectors of "static" frames, feature extractors, etc. Using this technique typical acquisition and preprocessing tasks can be easily uncoupled from the "consumer" applications.