AN AUTOFOCUS SYSTEM FOR VIDEO CAMERAS
AND DIGITAL IMAGE PROCESSORS
Thesis
Submitted to
Graduate Engineering & Research
School of Engineering
UNIVERSITY OF DAYTON
In Partial Fulfillment of the Requirements for
The Degree
Master of Science in Electrical Engineering
by
William Paul Blase
UNIVERSITY OF DAYTON
Dayton, Ohio
July 1988
© Copyright by
William Paul Blase
All Rights Reserved
1988
AN AUTOFOCUS SYSTEM FOR VIDEO CAMERAS AND DIGITAL IMAGE
PROCESSORS
APPROVED BY:
Donald L. Moon, Ph.D.
Advisory Committee, Chairman
Gary A. Thiele, Ph.D.
Associate Dean/Director
Graduate Engineering & Research
School of Engineering
Gordon A. Sargent, Ph.D.
Dean, School of Engineering
ii
ABSTRACT
AN AUTOFOCUS SYSTEM FOR VIDEO CAMERAS AND DIGITAL IMAGE
PROCESSORS
~
Name: Blase, William Paul
University of Dayton, 1988
Advisor: Dr. D. L. Moon
The development of and action of an algorithm for the automatic focusing of video
cameras, using digital image processing of video signal content, is described. This algorithm
examines the scene viewed by the video camera and measures the relative sharpness of edges
by calculating the two-dimensional first derivative, or gradient, of the image. The largest
peaks of the gradient, after appropriate filtering and thresholding, are added to get a "Focus
Merit Factor", which is indicative of the degree of focus of the image. The lens of the video
camera is then manipulated automatically to maximize the Focus Merit Factor, which results
in a focused image. Also described is a method for calculating and characterizing several
characteristics of the lens which affect the autofocus algorithm.
111
TABLE OF CONTENTS
APPROVAL PAGE
ABSTRACT
ii
. . . . . . . . . . . . . . . . . . .' . . . . . . . . . . . . . . . . . . . . . . . . . . . iii
TABLE OF CONTENTS
iv
~ LIST OF ILLUSTRATIONS
vi
LIST OF SYMBOLS
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ..
SPECIAL NOMENCLATURE.
ACKNOWLEDGEMENTS
INTRODUCTION.
viii
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ..
ix
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ..
xiii
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ..
1
CHAPTER
I.
AN INTRODUCTION TO AUTOMATIC FOCUSING METHODS.
. . . . 2
Lenses and Focusing
Methods of Automatic Focusing
II.
A METHOD FOR FOCUSING USING SCENE CONTENT . . . . . . ..
10
Introduction
Correction of Lens Distortion
Two-Dimensional Median Filters
Two-Dimensional Gradient Filters
Focus Merit Factor Calculations
Point of Focus Calculations
III.
A METHOD FOR CHARACTERIZING LENSES
28
IV.
ALGORITHM EVALUATION
35
IV
v.
CONCLUSIONS AND RECOMMENDATIONS
43
Conclusions
Recommendations
APPENDIX.
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ..
BIBLIOGRAPHY
45
51
v
LIST OF ILLUSTRATIONS
1-1. Camera Focus Behavior,
. , , . . . . . . . . . . . . . . . . . . . . . . . . . . . . ..
1-2. Ranging Autofocus System
6
1-3. Behavior of an Image of an Edge Under Changes in Focus
1-4. lllustration of Gradient Image-Out-of-Focus
~
6
1-5.111ustration of Gradient Image-In-Focus
. . . . . . . . . . . . 7
. . . . . . . . . . . . . . . . . . . . . 8
. . . . . . . . . . . . . . . . . . . . . . . . 9
2-1. Typical Convolution Function . . . . . . . . . . . . . . . . . . . . . . . . . . . "
18
2-2. Single and 16-Frame Average Images . . . . . . . . . . . . . . . . . . . . . . ..
19
2-3.111ustration of Scale Distortion Correction, Uncorrected Image
. . . . . . . ..
20
2-4. lllustration of Scale Distortion Correction, Corrected Image . . . . . . . . . ..
21
2-5. Two-Dimensional Separable Median Filter -Action
21
2-6. lllustration of the Action of 2-D Median Filter
2-7. Two-Dimensional Gradient Filter
. . . . . . . . ..
. . . . . . . . . . . . . . ..
22, 23
'. . . . . . . . . . . . . . . . . . ..
24
2-8. lllustration of Gradient Filter Action . . . . . . . . . . . . . . . . . . . . . . . ..
25
2-9. Peaks Produced by 2-D Gradient Filter
. . . . . . . . . . . . . . . . . . . . . ..
26
2-10. Focus Merit Factor vs Lens Position
. . . . . . . . . . . . . . . . . . . . . . ..
27
2-11. FMF Hill Climbing Algorithm . . . . . . . . . . . . . . . . . . . . . . . . . . ..
27
3-1. Example of Lens Scale Distortion . . . . . . . . . . . . . . . . . . . . . . . . . ..
31
3-2. 50mm Canon® Lens . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ..
32
3-3. Lens Characterization Target.
33
. . . . . . . . . . . . . . . . . . . . . . . . . . . ..
3-4. Radial Line and Vanishing Point Calculations.
. . . . . . . . . . . . . . . . . ..
33
. . . . . . . . . . . . . . . . . . . . ..
34
3-6. Illustration of Dot Center Finding Method . . . . . . . . . . . . . . . . . . . . ..
34
4-1. Grid Target
38
3-5. Found Dot Centers and Vanishing Point.
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ..
4-2. Lens Position versus Focus Merit Factor Curve for Grid Target
vi
38
4-3. Results of Autofocus Algorithm for Grid Target.
. . . . . . . . . . . . . . . ..
39
4-4. Lens Characterization Target . . . . . . . . . . . . . . . . . . . . . . . . . . . . ..
39
4-5. Lens Position versus Focus Merit Factor for Lens Characterization Target..
40
4-6. Autofocus Algorithm Results for Lens Characterization Target . . . . . . . ..
40
4-7. Plaster Face Mold
. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . ..
41
4-8. Lens Position vs Focus Merit Factor for Face. . . . . . . . . . . . . . . . . . ..
41
4-9. Lens Position vs Focus Merit Factor for Face, with Error Bars . . . . . . . ..
42
A-I. Examples of Slope Calculations.
50
. . . . . . . . . . . . . . . . . . . . . . . . . ..
vii
LIST OF SYMBOLS
An entire image is refered to by a capital letter, usually Roman, may be Greek; e.g., A, Y,
0. This may be modified by a subscript, e.g., Ax and Ay' if there are several related
images.
The individual elements of an image, or pixels, are referred to by adding indices to the image
symbol; e.g., A(x,y).
Vll1
SPECIAL NOMENCLATURE
Autofocus: an abbreviation for Automatic Focusing
Blur Circle: see Point Spread Function.
Convolution: the convolution of two continuous functions!
(x) and g (x) is defined as
00
! ** g = J ! (u)g (x-u) duo
(Bracewell 1978, 25).
-00·
The convolution of two continuous two-dimensional functions!
defined as
00
! **g = J
-00
(x,y) and g (x,y) is
00
J!
(x',y')g (x-x',y-y') dx'dy'
(Bracewell 1978, 243).
-00
In operations with discrete functions, such as images and image-processing operators,
the above integrals are replaced by convolution sums. For an MxN pixel image I(x,y)
and an operator H(x,y) (usually smaller than MxN) the convolution sum is
M-1 N-1
L
G(x,y)=I**H=L
I(m,n)H(x-m,y-n)
m=O n=O
(Krukar 1987, 35).
In practice, the operator H is usually specified as square, 2L+ 1elements on a side,
and the equation is specified thus:
L
L
L
G(x,y)=I**H=L
I(m,n)H(x-m,y-n).
m=-Ln=-L
The center element is then H(O,O)with H(-L,-L) in the lower-left corner. For instance,
given the operator
ix
H=
1
0
-1
1
0
-1
1
0
-1
4
4
4
4
4
4
4
5
5
5
5
5
5
5
6
6
6
6
6
6
6
6
6
(Le l )
and the image
1=
1
2
1
2
1
2
1
2
1
2
1
2
1
2
3
3
3
3
3
3
3
6
6
6
6
6
6
6
6
6
6
6
6
6
6
6
6
6
7
7
7
7
7
7
7
the result would be
G=
6
6
6
6
6. 6.
Note that the result Gis 2L pixels smaller (L on each side) in each dimension than I;
this is because the I, and thus the result of the convolution operation, is undefined for
all pixels less than a distance L from the edge. A second method to handle edge effects
is to extend I by L pixels on each side, using some arbitrary value (usually 0) in each
pixel, and act as if I was 2L pixels larger on a side than it really is. The result will
obviously be the size of the original image 1. Which method is picked depends upon the
operator chosen and the context in which it is used.
Correlation: In image processing literature, what is occasionally referred to as convolution is,
in fact, correlation - which is closely related and uses the equation
L
L
G(x,y)=I*H=L,
L, I(m,n)H(m+x,n+y).
m=-Ln=-L
The result of this operation is identical to that of convolution if the operator is reversed
left for
x
right and top for bottom. The reason that this is done is that correlation is much more
intuitive for a visual science like image processing; one simply overlays the operator
onto the image, with the center element H(O, 0) on top of the pixel being operated upon,
multiplies the two and adds the results. It also makes the computer code somewhat
easier to write and understand.
Depth of Field: The distance from the Focal Plane within which a target can still be said to be
in focus and a measure of how rapidly the FMF falls from its peak at the point of focus.
The Depth of Field is related to several physical characteristics of the lens; primarily,
however, it is inversely proportional to the aperture of the lens, usually set by the iris.
A pinhole lens with a very small aperture theoretically has an infmite depth of field and
never needs to be focused. For a lens with a large aperture, on the other hand, the
image of a target will blur very rapidly as it moves away from the focal plane.
Focal Plane: The plane in space, at some distance d from the focal point of the lens where d
depends upon the lens setting, at which objects form focused images on the image
plane. See Figure l-La.
Focal Point: The point in space behind a lens through which a11light rays entering the lens
parallel to the center axis of the lens will be bent to pass. The position of this point will
depend upon the lens setting. See Figure 1-1a.
Focal Stop (Near/Far): the mechanical stop(s) past which the lens focus ring may not be
turned. Also refers to the nearest and furthest (respectively) distances that the focal
plane may be from the lens.
Focus Measure Factor: see Focus Merit Factor.
Focus Merit Factor (FMF): A measure which indicates the degree of focus of an image.
D suaUy this takes the form of a Gaussian function of the distance that the focal plane is
from the target being focused upon and is at a maximum when the image is in focus.
See Figure l-Lb,
Xl
Focus Ring: The part of the lens which the operator turns (or otherwise adjusts) in order to
move the glass lens elements and focus the lens. See Figure 3-2.
Image Plane: The plane in space at which the image projected by the lens is formed. This
usually corresponds to the film or video sensor plane in a camera. See Figure l-la.
Machine Inspection System: any system which automatically inspects an item, usually
through the use of computers, video cameras, and image processors. See also Machine
Vision System.
Machine Vision System: a system wherein a computer uses one or several video cameras to
visually understand or inspect a scene or a target.
Point Spread Function: a two-dimensional Gaussian function (also known as a Blur Circle)
which is convolved with a perfectly focused image to obtain a blurred image. This
function is dependent on the particular lens in use.
Xli
ACKNOWLEDGEMENT
I thank: the members of the Advanced Systems Research Group, AFW ALl AAA T-3, for their
support, especially Mr Richard Jones, Chahira Hopper, and Dr Louis Tamburino.
xiii
INTRODUCTION
An automatic method for focusing video cameras is useful, or even vital, in cases
such as machine inspection systems, scanning electron microscopes, and electro-optic sensor
system where the operator either cannot manually focus the camera satisfactorily or is too
busy doing other things. Previous to the advent of digital image processing systems,
autofocus systems had to rely upon mechanical triangulation or range-finding mechanisms,
which have inherent shortcomings in range and/or accuracy and which add additional
equipment onto the camera assembly. With the use of image processors, especially if they are
being used in the system to which the camera is attached anyway, an autofocus system can
use the video image itself to calculate the point of optimum focus for the lens. Additional
advantages come from the fact that a single focusing algorithm can be used for many
different types of systems with no regard to the exact type of hardware that is attached to the
video camera.
The autofocus system described in this paper was developed by the author at the
Avionics Laboratory, Wright-Patterson Air Force Base, Ohio as part of the work on the
Numerical Stereo Camera, a machine inspection system being developed by the Advanced
Systems Research Group.
1
CHAPTER!
AN INTRODUCTION TO AUTOMATIC FOCUSING METHODS
Lenses and Focusing
Refering to Figure I-la, a camera image is in focus when light rays reflected from
some target at a distance dt (thefocal plane) from thefocal point produce a sharp
(unblurred) image on the image plane of the camera (usually the film or sensor plane). If one
considers the target object and the camera to be stationary, the distance d1 from the focal
point to the image plane and the distance d2 from the focal point to the focal plane form a
fixed ratio. As the focus ring on the lens is adjusted, the focal point moves further from the
image plane and thus the focal plane will move from quite close to the lens (the near focal
stop) to quite far from the lens (thefar focal stop, usually labeled as 00). The further the
target is from the focal plane, the more blurred the image will be.
A blurred image can be described as a perfect image that has been convolved with a
two-dimensional Gaussian function, also called apoint spreadfunction
or blur circle
(Shazeer and Harris 1985, p150), with a height of unity. The further the target is from the
focal plane and the more out of focus it is, the wider the function is. Obviously, when the
image is perfectly in focus the point spread function has no width at all (a Dirac
or impulse
function with an area of unity). A Focus Measure Factor, or Focus Merit Factor, (FMF) is a
measure of the degree of focus of an image. Usually the FMF is at a maximum when the
image is in focus and is inversely proportional to the width of the point spread function,
depending upon the particular algorithm in use for computing it. Figure 1-1b shows the
behavior (typical) of the FMF according to the distance from the focal plane to the target. The
exact shape of the curve obtained depends upon the method used for computing the FMF and
upon the characteristics of the target, primarily the number and type of edges, although it is
usually Gaussian in nature.
2
3
Methods of Automatic Focusing
There are two primary methods used for automatically focusing the lens of a video
camera: ranging
and scene-based-focusing . Each method has its own advantages and
disadvantages and which one is chosen will depend upon the system in which the autofocus
system will be used.
Ranging
In a ranging system, the distance from the camera to the target is found by some
means and the lens, which must be pre-calibrated, is adjusted accordingly. Figure 1-2 shows
the two basic types ofranging systems: the triangulation
system and the time-a/flight
system. Ranging systems have the advantages of being simple and inexpensive. Infra-red
triangulation autofocus systems are found on many consumer video and 35mm cameras. The
Polaroid® Ultrasonic Ranging System® time-of-flight'system is widely used both in
Polaroid® cameras and other commercial distance measurement systems. Ranging systems,
however, have three disadvantages. First, separate apertures and/or mounts are required for
the ranging sensors; the triangulation approach requires two apertures which must be widely
separated for accuracy. Secondly, anything which interferes with the sensing beam, such as
a piece of glass, will negate the autofocus system. Thirdly, neither approach is easily capable
of focusing on an arbitrary area of the target, other than at the center of the image.
Scene Based Focusing
Scene based focusing systems overcome these disadvantages by adjusting the lens
according to some characteristic of the image itself (or some arbitrary subsection), usually
through some Focus Merit Factor (FMF) which is derived from the scene and maximized (as
described previously). The two scene based systems which have been most reported in the
literature derive the FMF either according to the high frequency content of the image (see
Hanma, et al. 1983; Schlag, et al. 1983; and Kokoulin and Poleshchuk 1978) or according
4
to the height of the peaks caused by edges in the two-dimensional first derivative (gradient)
of the image (see Shazeer and Harris 1985; Tee, Smith, and Holburn 1979; and Northrop
1984). Both approaches essentially measure the sharpness of edges in an image; the second
by a more direct method than the first.
Obtaining the Focus Merit Factor from the high-spatial frequency content, as
measured by a two-dimensional Fourier Transform, derives from the fact that the frequency
spectrum of a sharp edge contains many high frequency harmonics. The pictures in Figure
1-3, recorded off of the video display of the Gould image processor, shows an edge formed
by a piece of white cardboard covering one-half of a piece of black cardboard. Superimposed
upon the bottom of each picture is a graph which shows the intensity (y
axis) of each pixel
along a cut through the center of the image - indicated by the horizontal line across the center
of the picture. Figure 1-3a shows the edge when the lens is as out of focus as possible, 1-3b
shows the same edge when the lens is sharply focused. As described above, the point spread
function causes the edge to spread out and become rounded as the image goes out of focus,
which results in a lessening of the high frequency harmonics in proportion to the degree of
defocus. The Focus Merit Factor, then, would be made proportional to the area of the spatial
frequency spectrum.
There are two problems with this approach. The first is that the two-dimensional
FFT (Fast Fourier Transform) of an image takes a considerable amount of time to calculate.
The second occurs if the image has large amounts of homogeneous (blank) area, or if there
are no sharp edges. The high spatial frequencies in the two-dimensional Fourier Transform
caused by noise in the video system - which will remain constant regardless of the lens
position - may then overwhelm any high frequencies that are caused by edges.
The answer to these problems, as proposed by Shazeer and Harris (Shazeer and
Harris 1985) and expanded on in this thesis, is to ignore anything that does not contribute to
the focusing operation and concentrate on the edges (or other steep gradients).
We have found that almost all focus information is found at edge discontinuities
and almost no information is contained in other scene regions.
One of the problems of focusing schemes that evaluate the high frequency
content of a scene is that they take into account large portions of the image that
are almost totally devoid of information but still contain noise, some of which is
in the information bandwidth. This algorithm is fundamentally different from
others because it is spatially selective and looks only at the edges where the
information and the corresponding SNR [Signal to Noise Ratio] are maximum
(Shazeer and Harris 1985, 150,151).
5
Figure 1-4a shows the same edge as Figure 1-3a with an image of the horizontal
gradient (two-dimensional first derivative) superimposed over the center one-half of the
picure. The intensity of the gradient image corresponds to the steepness of the slope
(compare the brightness of the gradient image overlay and the intensity of the corresponding
point in the image of Figure 1-3a) and an edge in an image results in a large peak in the
gradient. Figure 1-4b shows the same gradient image with the same type of cross-sectional
intensity graph as in Figure 1-3 superimposed. As the lens is moved into a focused position,
the peak will grow larger, reaching a maximum when the image is perfectly focused, as
shown in Figure 1-5, a and b. The FMF for this method is calculated from the relative
heights of peaks in the gradient image, generally by summing or averaging all of the peaks in
the gradient image after thresholding or filtering to remove lesser peaks which are due to
noise. There are three advantages to this approach. First, since it is the relative size of the
peaks that are important, filtering techniques may be used which reduce noise at the expense
of high frequency content. Second, since only the edges and other steep gradients are used in
the calculation of the Focus Merit Factor, large, homogeneous spaces, and thus a great deal
of noise and irrelevant material, are ignored. Third, if proper filtering is used to remove noise
then there need be no sharp edges in the scene at all, just enough of an intensity gradient
somewhere to produce a peak in the gradient image.
The rest of this thesis will report on the scene based, gradient-peak detecting
autofocus system developed in the Advanced System Research Group laboratory. Chapter II
will present the details of the algorithm, Chapter III will present a method for measuring
characteristics of lenses which effect the operation of the autofocus algorithm, Chapter IV
will present experimental results, and Chapter V will offer suggestions for further
improvements and research.
6
IMAGE PLANE
~
LENS
a
•....•
~
C'I
~
•....•
~
0
0
('f"l
ll.
ll.
0
~
~
u u u
~
~
~
~
~
en
FOCAL PLANE 2
FOCAL PLANE 3
~
0
E-<
E-<
•....•
FOCAL PLANE 1
ll.
f2 f2 f2
b
;:J
U
f2
DISTANCE
Figure 1-1: Camera Focus Behavior
a
I
CMffiRA ~~N!::S===========B=EAM====~~
TRANSCEIVER
Figurel-2: Ranging Autofocus Systems
b
7
a - Out of Focus
b - In Focus
Figure 1-3: Behavior of an Image of an Edge Under Changes in Focus
8
a- Gradient Image of Out-of-Focus Edge
b - Gradient Image of Out-of-Focus Edge, With Profile
Figure 1-4: Illustration of Gradient Image - Out-of-Focus
,
-
----
i
9
a- Gradient Image of In-Focus Edge
b - Gradient Image of In-Focus Edge, With Profile
Figure 1-5: illustration of Gradient Image - In-Focus
CHAPTERll
A METHOD FOR FOCUSING USING SCENE CONTENT
Introduction
Shazeer and Harris, in their paper reporting work at Honeywell on FUR (Forward
Looking Infra-Red) imaging systems (Shazeer and Harris 1985), proposed a scene based,
gradient-peak detecting, automatic focusing algorithm consisting of three steps: (1)
convolving each scan line F(x) in an image with a digital filter mask H(x), of length 2N+ 1,
producing an approximation to the first derivative F'(x):
N
F'(x) = ~ H(i) F(x-i);
i=-N
(1)
(2) rectifying the result to produce a positive-only result; and (3) forming the Focus Merit
Factor as the average value of the peaks (a peak is defined as a pixel which is greater in
intensity than the pixels on either side, out to some arbitrary distance) which exceed a
predetermined threshold (see Figure 2-1). The digital filter mask serves two purposes: it
calculates the gradient of an image at some pixel A(i,j) and it performs some noise
suppression by also using the values of surrounding pixels in an averaging process. These
three steps would be repeated on consecutive frames, while some mechanism adjusted the
optics so as to maximize the Focus Merit Factor and arrive at the sharpest image.
As an example of this algorithm, suppose that one was using the filter
H(x)={ 1,1,0,-1,-1 }(N=2) on a scan line of length M=16 with pixels having the following
intensity values:
°° °
5
°
x=15
2
7
10
10
10
8
3
°°°°°
(2)
11
which might represent a white stripe (values of 10), perpendicular to the scan line, with
sloped edges on a black background (values of 0) and a noise spike in pixel 3. The result,
after the filter operation, and using the first convention, would be:
o
0
5
2
4 12 18
9
-6 -17 -18 -11 -3 0
0
o.
(3)
2
4 12 18
9
6 17 18 11
0
O.
(4)
Rectified, this would be:
o
0
5
3 0
Defining a peak as a pixel that is greater than its four nearest neighbors (two on either side),
there are peaks in pixels 2 (=5),6 (=18), and 10 (=18). Assuming that, due to prior
calculations, the threshold has been set at 10, the FMF would be 18+ 18=36.
The exact shape of the filter and the level of the threshold are determined through
statistical methods according to the noise present in the system. The paper by Shazeer and
Harris shows the exact methods for doing so, which are too complex to address here.
Since the Shazeer-Harris algorithm described above was designed for use in a
system that had severe size and processing restraints (e.g. a FUR pod on a fighter aircraft)
the algorithm is essentially a one-dimensional one, that is it operates horizontally on pixels
along a scan line, but not vertically between scan lines. This limitation means that it will
not take advantage of edges in the image that run parallel (or nearly so) to the scan lines of
the image. Neither does it take advantage of additional filtering that could reduce the effects
of noise. The method proposed in this thesis, therefore, is as follows: 1) correct the image
for certain distortions caused by the camera lens; 2) filter the image with a two-dimensional
median filter to reduce the shot noise produced by the image intensifiers in the video cameras;
3) calculate the two-dimensional gradient; and 4) sum the peaks, after thresholding, to form
the Focus Merit Factor. The remainder of this chapter will explain the details of the algorithm
and methods used for calculating the threshold and for fmding the maximum FMF.
Chapter 1 pointed out that scene based autofocus methods have an advantage over
ranging methods in that they can focus on various portions of the image; this is particularly
useful when the target to be focused is in front of a background which is fairly far from the
camera.This autofocus algorithm is capable of operating on any portion (usually rectangular,
for simplicity) of the original image designated by the operator. In the descriptions which
12
follow, all of the various algorithms apply equally well to a portion of the image as to the
whole.
Note: the images presented as input to the algorithm ('A' in the equations below) are
in fact averages of between 2 and 8 video frames. Although the noise does have a non zero
mean, meaning that averaging over any number of frames cannot remove it, averaging does
reduce the noise a significant amount. Figure 2-2 shows a single frame image (a) and an
image which is the average of 16 video frames, the number usually used in the lab (b). Upon
close inspection, one can see that a considerable amount of 'speckle' type noise is removed
from the image by the averaging process.
Correction of Lens Distortion
Most camera lenses introduce a 'zoom' or scale distortion into the image due to the
movement of the lens elements during the focusing adjustment. Figure 2-3 illustrates this
distortion by showing a test grid (one meter from the lens) with the lens set to focus at 0.9
meter (a) and
00
(b); overlaid on the video image is a rectangular box formed by the cursors
of the image processor. In Figure 2-3a the cursor box encloses an area of the test grid
approximately 7 by 5.5 squares. In Figure 2-3b the cursor box, which was not changed
between the two photographs, now encloses an area 8.5 by 6.5 squares. The lines of the test
grid are straight and orthogonal, the geometric warping in the image is caused by the video
tube used in the camera. Because this warping does not depend upon the lens position it is
not corrected for in this algorithm.
Each pixel A(i,j) in the video image is moved from its position in the 'ideal' image
A'(i,j) (the image which would be seen with an ideal lens and camera and which does not
really exist) along a radius centered on a 'vanishing point' near the center of the image. The
pixel is moved by some factor Alpha, which is a function of the current position of the
focusing ring of the lens. Generally, the image is largest when the lens is at the closest focus
stop (usually on the order of 0.5 meter) and smallest when the lens is at the other stop
(usually designated
00
).By convention, the function which calculates Alpha is calibrated
when the lens is characterized so that Alpha is equal to 1 (and thus A = A') when the image
is largest. For other lens positions, as the focusing ring is moved towards the
00
setting, A is
smaller than A' and Alpha is greater than 1 (usually between 1 and 1.2) - describing the
amount that A must be enlarged to correspond to A'. Figure 2-4 shows the result of applying
13
this correction to the portion of Figure 2-3b within the cursor box; the corrected area is
within the rectangle in the center of the image. Chapter 3 explains this scale distortion in
further detail as well as the algorithm used to characterize the lens.
Two-Dimensional Median Filters
The particular filter used for preliminary noise reduction is a recursive,
two-dimensional
separable median filter (Arce, Gallagher, and Nodes 1986, 116 and
following). In general, a filter takes an input signal A(x) and produces some output signal
Y(x). In digital signal processing (of which image processing is a two-dimensional example)
the signal is sampled at discrete intervals and the output signal y is calculated by convolving
the transfer function of the filter with the input signal, as demonstrated in the explanation of
the Shazeer and Harris algorithm above. In general, each data point in the output signal Y(x)
corresponds to a particular point A(x) in the input signal and is calculated through some
function F(z) of that point and the surrounding points:
Y(x)=F(m)A(x-m)+F(m+l)A(x-m+l)+
...+F(-I)A(x-l)+F(O)
A(x)+F(I)A(x+l)+
+F(n)A(x+n).
...
(5)
Please refer to Bracewell (1978) or any text on digital signal processing for a more complete
explanation of the process. In a median filter, Y(x) is the median of 2N+ 1 data points evenly
centered on x, N points on either side (N is always odd, to ensure a unique median value).
Using the same string of data used previously, and proceeding from left to right,
x=O
o 0
0
5
0
2
x=15
7 10 10 8 3 0 0 0 0 0,
(6)
a median filter of width 3 (N=I) would produce Y(x) equal to
o 0
0
0
2
2
7
10 10 8 3 0 0 0 0 O.
(7)
Pixel 4 (=2), therefore, in Y(x) is the median of Pixels 3, 4, 5 (=5, O,and 2). The first and
last N values are artificial and arbitrary, since the image is not defmed past its edges. When
the window is centered on the first or last points in the data stream Arce, et al. recommend
14
filling in the blanks with that data point (0 in this case).
Quoting Arce, et al., some additional concepts:
A constant neighborhood is a region of at least N+ 1 consecutive points, all of
which are identically valued.
An edge is a monotonically rising or falling set of points surrounded on both sides
by constant neighborhoods.
An impulse is a set of N or fewer points whose values are different from the
surrounding regions and whose surrounding regions are identically valued constant
neighborhoods.
A root is a signal that is not modified by filtering (Arce, Gallagher, and Nodes
1986,93).
Generally the root signal, which consists of only constant neighborhoods and edges,
is the desired end result of the filtering process; the result of the previous example is a root
signal. The filter described in the example above was non-recursive,
that is there was no
feedback of already processed signal values into the input stream. A recursive median filter
uses the points Y(x), Y(x-1), ...,Y(x-N) instead of A(x),A(x-1), ...,A(x-N). It can be shown
(Arce, Gallagher, and Nodes 1986,97) that a recursive' median filter arrives at a root signal
in one pass, instead of the several possibly required by a non-recursive filter, and that it has a
greater smoothing effect
The particular property of median filters that we are interested in here is that any
impulses of width N or less are completely filtered out of the signal, while plateaus, or
constant neighborhoods of N+ 1 or greater points, are left alone. Again given the signal A(x)
x=O
x=15
000502710108300000
(8)
A recursive median filter of width 3 (N=l) would produce Y(x) =
o
0
0
0
0
2
7 10 10 8 3 0
0 0 0 O.
(9)
A two-dimensional separable median filter processes an image in two passes, as is
illustrated bu Figure 2-5. In the first pass (a), a horizontally oriented median filter is passed
over the rows, working from left to right, top to bottom; in the second (b), a vertically
15
oriented filter is passed over the output of the first, Please refer to Arce, Gallagher, and
Nodes (1986) for a full explanation of the properties of this kind of filter.
In our experience, the noise produced by the image intensifiers in the cameras used
in the AAA T-3 laboratory is shot noise, producing spikes one or two pixels in width by one
pixel in height. If not removed, these spikes would produce large, anomalous peaks in the
gradient image which might overwhelm legitimate peaks (note that the average height of the
peaks caused by the noise would not change as the lens is moved since the noise is
introduced into the system after the lens). Since most lines or other edge producing features
of an image are greater than three pixels wide a recursive separable median filter with N=1
removes most of these spikes very nicely without altering any of the other characteristics of
the image necessary for the gradiant production process. A filter with N=3 may also be used,
depending upon the fmeness of the detail in the image (any lines less than N pixels in width
will also be removed). Figure 2-6 shows an example of the operation of various median
filters on an image.
Two-Dimensional Gradient Filters
After the original image A has been cleaned up with the recursive separable median
filter, the gradient images Gx (the gradient in the x (row) dimension) and Gy (the gradient in
the y (column) dimension) are produced by convolving the result, Y, with a two dimensional
gradient filter, Actually, a full gradient image G would contain both the magnitude (GM(x,y)
= ~( Gx(x,y)2+ Gy(x,y)2 ) ) of the gradient and the angle ( G0(x,y)=arctan
Gy(x,y)/Gx(x,y)
), but all we are interested in here is the magnitude of the individual results.
As will be shown later, the Focus Merit Factors are computed seperately for each dimension,
to save computation time, and then combined. Figure 2-7b shows the gradient filter used, it
is a variation of the Sobel edge detector (Ballard and Brown 1982, 77) and averages the
slopes of a number of lines through pixels around the center pixel. Strictly speaking, the
filter should be as shown in Figure 2-7a, but since this algorithm is only interested in the
relative values of the gradient, the entire filter (and thus the result) is multiplied by a
constant to make all of the elements integer values, making the calculations faster. The same
filter is used in both the x and y dimensions, being rotated 90 degrees counterclockwise.
Other sizes of filters are possible; the 5x5 element filter was chosen to balance between noise
16
reduction and overmuch smoothing of data (larger filters smooth more, so fine detail in the
original image could be lost, smaller filters don't get rid of noise as well). Please refer to the
Appendix for further information on the derivation of these filters. Figure 2-8 shows
examples of the output from the 5x5 gradient filters (Horizontal and Vertical), operating on
the image of Figure 2-4. Note how the bright stripes in Figure 2-8 correspond to the edges of
the lines in Figure 2-4.
Focus Merit Factor Calculations
Once the gradient image for a dimension has been generated, the Focus Merit Factor,
or FMF, may be arrived at by calculating the average height of any peaks that exceed some
threshold t . Because long edges in the image form ridges ill the gradient image, and a point
on a ridge is usually a maximum only perpendicular to the ridge, the algorithm does not look
for points that are peaks in two dimensions. Instead, an FMF is calculated seperately for each
dimension (FMFx and FMFy) and the two are combined at the end for a total FMF; FMFc =
-V(FMFx2+FMFy 2), where FMFx is derived from Gx and FMFy from Gy. This is derived
from the fact that, as pointed out above, the gradient at any point in the image is the vector
sum of the x and y gradient vectors for that point. Calculating the Focus Merit Factors
separately assures that the algorithm will focus well on an image no matter which direction
the edges run.It also assures that a large number of peaks will be considered in the
calculations, which lessens the impact of noise or other errors on anyone peak.
A point G(x,y) is considered a peak if it is greater than its two neighbors on one
side, G(x-2,y),G(x-1,y)
for the x dimension (rows); G(x,y-2), G(x,y-1) for the y dimension
(columns), and greater than or equal to its neighbors on the other, G(x+ 1,y) and G(x+2,y)
or G(x,y+ 1) and G(x,y+2) respectively. The 'greater than or equal to' is arbitrary and merely
assures that if the algorithm encounters a plateau of constant values that one edge of the
plateau will be counted as a peak. The choice of two neighbors on either side is arbitrary and
is based on observations that most ridges produced by lines and edges in the original image
are three or four pixels wide. Figure 2-9 shows the peaks that were found in the gradient
images of Figure 2-8.
The thresholdt is calculated as the average height of peaks due to 'edges' in the
signal caused by noise. Because the shot noise, both from the intensifiers and from the video
17
electronics, is signal dependant (Kasturi and Walkup 1986), merely measuring the noise
with the lens cap on is not satisfactory. In addition, we want to take into account the effect of
the various filtering operations on the noise leveLThe solution that we have found is to take
the image D which is the absolute value of the difference of two consecutive images. This
removes the scene, which doesn't change between images, and leaves the noise, which does.
D is then processed through the median fliter and gradient filter to produce an image which
contains only peaks due to noise present in the original image(s). The threshold t is then then
calculated in the same way as the FMF, as described above. Since the noise is not
particularly oriented in anyone direction the same threshold is used for both x and y
dimensions.
Point of Focus Calculations
Figure 2-10 shows a graph of the FMF vs the lens position for the image shown in
the previous figures. The bell shaped nature of the curve and the peak at the point of best
focus can be clearly seen. The method used to calculate the peak of the Focus Merit Factor
curve, where the image is in focus, is one of simple hill climbing. Refering to Figure 2-11,
the algorithm begins by gathering 5 FMF points to find out which way is 'uphill', moving
after each point a distance ~ from the previous one towards the lens stop that is the farthest
from PO. ~ is arbitrarily calculated as 1/50 of the total lens travel range. After the initial 5
points are gathered, a line is fitted to the points using the least-squares method and the slope
of the line is noted. The algorithm then proceeds to gather additional points, traveling uphill
from PO. After each point is gathered a line is fitted to that point and the previous 4, and the
slope is again examined. At some point after the algorithm passes over the peak of the FMF
curve the slope of the fitted lines will change sign. When this occurs, the algorithm gathers 5
more points, so as to get a number of points on both sides of the peak for more accurate
curve fitting, and fits a second order polynomial to the last 20 data points gathered. The
curve is fairly close in shape to a parabola at the top of the curve, and so the peak of the fitted
polynomial, where the first derivative is equal to 0, is declared to be the peak of the FMF
curve and the point of optimum focus.
18
NUMERICAL PIXEL VALUE OF
LINE OF VIDEO DATA
r-r-...,.-......,. .....•..r
f (t\
1
IN FOCUS
~
OUT OF FOCUS
t (DISTANCE ALONG
SCAN LINE)
L
I
IDEAL FILTER SHAPE
FILTERMASK(TOSCALE)
L::. J
t------....,.~to::~-~~--..,.......,,,,....--
.••••
V
NOTCH
+- ACTIJALDIGITAL
FILTER SHAPE
5 PIXELS
INFORMATION REGION
FOCUSED SCENE
f rt\
2
f2(t) IS THE CONVOLUTION OF fl (t)
wrrn THE FILTER MASK (APPROXIMATES
THE FIRST DERIVATIVE)
Figure 2-1: Typical Convolution Function (Shazeer and Harris 1985, 153)
19
a - Single Frame
b - 16 Frame Average
Figure 2-2: Single Frame and 16-Frame Average Images
20
a - Lens at Near Stop (.9 m)
b - Lens at Far Stop (00)
Figure 2-3: Illustration of Scale Distortion Correction, Uncorrected Image
21
Figure 2-4: illustration of Scale Distortion Correction, Corrected Image
p~.
"".
•
ORIGINAL PIXELS
RESULTS OF FIRST PASS
ELEMENT BEING
FILlER WINDOW
~
MODIFIED
~
e RESULTS OF SECOND PASS
'-.
e e 01~ e 0 •••••••••••••
-_
••. DIRECTION OF WINDOW MOVEMENT
•••••••••••••••••••
•••••••••••••••••••
•••••••••••••••••••
•••••••••••••••••••
a
eo@ee@oeeo@ee@oeeo@
0000@00@0000@00@@00
@0~~0~~0~~0@~0@~~0~
~00~~0~~0@@0@~00~00
~~00~00~e0~G0~@0@~0
Figure 2-5: Two-Dimensional Seperable Median Filter-Action
b
b - Result of 3x3 Filter (Ne l )
Figure 2-6: illustration of the Action of 2-D Separable Median Filter
23
c - Result of 5x5 Filter (N=3)
d - Result of 7x7 Filter (N=5)
::.=:::re 2-6: Illustration of the Action of 2-D Seperable Median Filter (Continued)
24
-0.025
-0.050
0.000
0.050
0.025
0.025
0.025
0.025
0.025
0.025
-0.025
-0.050
0.000
0.050
0.025
0.050
0.050
0.050
0.050
0.050
-0.025
-0.050
0.000
0.050
0.025
0.000
0.000
0.000
0.000
0.000
-0.025
-0.050
0.000
0.050
0.025
-0.050
-0.050
-0.050
-0.050
-0.050
-0.025
-0.050
0.000
0.050
0.025
-0.025
-0.025
-0.025
-0.025
-0.025
Vertical
Horizontal
a - 5x5 Gradient Filter
-l.
-2.
O.
2.
l.
l.
l.
l.
l.
l.
-l.
-2.
O.
2.
l.
2.
2.
2.
2.
2.
-l.
-2.
O.
2.
l.
O.
O.
O.
O.
O.
-l.
-2.
O.
2.
l.
-2.
-2.
-2.
-2.
-2.
-l.
-2.
O.
2.
l.
-l.
-l.
-l.
-l.
-l.
Vertical
Horizontal
b - Scaled Gradient Filter
Figure 2-7: Two-Dimensional Gradient Filter
25
a - Horizontal Filter
b - Vertical Filter
Figure 2-8: illustration of Gradient Filter Action
26
b - Vertical
Figure 2-9: Peaks Produced by 2-D Gradient Filter
27
700~----------------------------------------------------~
600
500
~
~
~
400
300
200~~TT~~~TTMM~~TrMM~~TrnT~~~~rn~~~~~MM~
234
o
1
5
6
7
Lens Position (Volts on Display)
Figure 2-10: Focus Merit Factor vs Lens Position
7~r---------------------------------------------------~
Parabola:
y
•
= - 734.98 + 565.07x - 57.61x"2
Preliminary points
Points included in
parabola
•••
1
3
••
•
•
PO·
o
2
4
5
Lens Position (Volts on Display)
Figure 2-11: FMF Hill Climbing Algorithm
6
7
CHAPTER III
A METHOD FOR CHARACTERIZING LENSES
In any system which relies upon accurate control of a lens for operation, knowledge
of that lens's characteristics is essential. Any lens will warp or change the image presented to
the video camera in various ways, some of which will affect the results of whatever
algorithm is the purpose behind the system, and some of which will have negligible affect.
A lens will also have certain characteristics which, while they may not affect the results, the
knowledge of which may assist the operator. In general, there are three such characteristics
for any given lens. First, the lens may warp or distort the image in a way that is constant
under all conditions of use. Barrel distortion and astigmatism due to imperfections in the
\
shapes of the glass elements in the lens are two such distortions.Secondly, the lens may
distort the image in a way that changes with the focus setting of the lens. In the Canon®
lenses used in the AAA T-3 laboratory, the movement of the lens elements as the lens is
adjusted from one stop to the other results in a 'zoom', or expansion/contraction
affect,
illustrated in Figure 3-1. Each point in the image moves along a radius toward or from a
'vanishing point', with the change in radius being proportional to the amount that the focus
ring was adjusted. Finally, because the position of the lens is read electronically by the
autofocus mechanism, there will be a relationship between that reading and the distance to
which the lens is currently focused, which can usually be read on a scale inscribed on the
focus ring of the lens. The lens position is read electronically as a voltage across the feedback
potentiometer, which is directly proportional to the focus ring position, while the current
focal distance is shown on the focus ring. A typical lens, in this case a 5Omm, is shown in
Figure 3-2. For clarity of illustration this lens is lacking the focus motor and feedback
potentiometer described above.
In the autofocus system described in the previous chapter, any warping or distortion
introduced by the lens which is constant over the focus range does not affect the algorithm
28
29
significantly. Therefore, this, as well as the distortions introduced by the video camera tube,
are not corrected for. The second kind of distortion, the 'zoom' affect, however, does
introduce distortions and must be corrected for. Since the algorithm detects and measures the
relative sharpness of edges within an area of the image, it will be disturbed if edges 'enter' or
'leave' the area as the lens is adjusted towards optimal focus. The third lens characteristic,
the relationship between the lens' focal distance setting and the electronic reading, is not
essential to the algorithm's operation, but is useful in checking results. The relationship
between the two is linear and is simply calculated.
The 'zoom', or scale, distortion is characterized by the use of the target shown in
Figure 3-3. As illustrated in Figure 3-4, the center of each of the four large dots is found for
a number of positions of the lens over the focus range and the radial line that each appears to
travel 'Overis calculated. The intersection of the four lines then will be the 'vanishing point',
or the center of the scale distortion - noted as Xc and Yc- Alpha, the measure of the scale
distortion, is then calculated for each lens position by averaging the radii from the vanishing
point to each dot center and dividing this into the average radius for the lens setting where the
image is largest - usually the setting with the closest focal distance. Since Alpha is a linear
function of the lens setting, a line is fitted to the found values, using the least-squares
method, and the coefficients Mc and Bc are found and saved. This allows the autofocus
program to calculate Alpha for any lens position. Because the center of the dots cannot be
found exactly and because the scale change may not be precisely determined, the found
positions of each dot center do not form a perfect line, nor do the four lines all meet at the
same point. Therefore, the algorithm uses the method of least squares to fit a line through the
found centers of each dot. These lines will form a rough square at the center of the image,
surrounding the vanishing point, which is calculated by finding the center of the square.
Figure 3-5 shows the result of an actual application of the procedure on the 85mm lens: the
found centers of each dot for each lens setting and the 'vanishing point' are superimposed
over an image of the target taken with the lens at the O.9m setting - where the image is
largest.
The dots on the target are large (1/2 inch in diameter) so that they can be
distinguished over the entire focus range of the lens - during a large portion of which they
will be out of focus. The center of each dot is found by the following procedure, illustrated
in Figure 3-6. First, the operator, using a joystick, places a cursor within the dot as seen on
the screen of the image processor. Knowing that the cursor is within the dot, although
,
30
probably not near the actual center, the program then draws eight radial lines from the cursor
location. Each radial is 30 pixels long, since we are using an optimised target and the dot is
known to be less than 60 pixels in diameter and is at a multiple of 45 degrees
(0,45,90, 135,etc). The first derivative of each line (the derivative as defmed in the direction
away from the dot center) is then calculated. Since the dots are black on a white
background, the dot border will produce a very strong peak in the derivative, even if the dot
is extremely blurred. This peak is then found along each radial and the center of the circle
formed by the eight peaks is declared to be the center of the dot. It should be noted here that
by the design of the target, the edge of the dot is the strongest dark-to-light transition in the
vicinity of the dot center; even if the radial line goes off of the edge of the target, this edge
will be light-to-dark and will produce a negative derivative peak.
The electronic reading/focus position relationship is found by having the operator
use the settings as marked on the focus ring as positions during the above algorithm and
entering both the feedback voltage reading and the focus ring setting. The two follow the
relationship
(10)
where
D is the setting on the focus ring
V is the feedback voltage
and
Mf and Bf are linear coefficients.
For the 85mm lens used in the examples, MF1.64 and BFO.73, with D ranging between
10m and 0.9m and V ranging between 0.60 and 5.95 volts.
31
a - Target Grid at Far Focus Stop (00)
b - Target Grid at Near Focus Stop (.9 m)
Figure 3-1: Example of Lens Scale Distortion
32
(
a - Lens at 00 Setting
b - Lens at .6m Setting
Figure 3-2: 50mm Canon ® Lens
33
Figure 3-3: Lens Characterization Target
FOUND DOT CENTERS
"I
~-
FOUND CENTER OF
EXPANSION
C4
C3
B3
FITIED LINE?
Figure 3-4: Radial Line and Vanishing Point Calculations
34
(
Figure 3-5: Found Dot Centers and Vanishing Point
A
D'
C'
~~
'~.
••
~
B
~/
INITIAL CURSOR LOCATION
~_C
FOUND CENTER OF DOT
f'D
DOT
DERIV ATIVE AWAY FROM CURSOR
FOUND EDGES OF DOT
II'
C'
C
~
~SOR
C' ~----~~----LOCATION ~
"FOUND
CENTER
Figure 3-6: illustration of Dot Center Finding Method
CHAPTER IV
ALGO~EVALUATION
(
Figures 4-1 through 4-10 show the operation of the algorithm described in this
thesis for three targets: the camera test grid from Chapter 2, the Lens Characterization Target
described in Chapter 3, and a plaster model of a human face. The targets are shown in order
of increasing difficulty of automatic focusing and the lens size is 85mm. Please note that in
Figure 4-2 and all of the graphs that follow, the lens position is given in Volts Displayed,
which is simply the voltage registered across the feedback potentiometer attached to the lens,
which in turn is inversely proportional to the square root of the lens position. As mentioned
in Chapter 3, the relationship between between the feedback voltage and the actual focal
distance is --JD = Mr(lN)+Bf
where D is the setting on .the focus ring (in meters), V is the
feedback voltage, and Mf and Bf are linear coefficients (in this case 1.636678 and .7271264
respectively); thus a lower voltage reading corresponds to a focal distance farther from the
lens. If the lens position were given in meters, the graph would be crunched to one side and
much more difficult to read.
It must be pointed out that there was no subjective test available as to whether the
image was actually focused optimally or not. The only reference was the lens position noted
when the image was focus manually via a video monitor. Focusing by eye, however, was
more than adequate for demonstrating that the theory behind the algorithm is valid and that
the peaks in the FMF curves were in the correct places. In every case tested the lens focused
where there was a peak in the FMF curve.
The test grid, shown in Figure 4-1, is nearly an ideal target. It has many closely
spaced lines running both horizontally and vertically and thus a great many edges to focus
upon. Figure 4-2 shows the Focus Merit Factor versus the Lens Position as calculated for the
area within the cursor box overlayed on the preceeding photograph over the entire focal range
of the lens (O.9m to 00). The target was approximately one meter from the lens and the lens
35
,
36
position display indicated approximately 5 Volts when the lens was focused by eye from the
video monitor.
Figure 4-3 shows the action of the Autofocus Algorithm on the test grid target. The
algorithm was started with the lens at approximately the 1.9 meter setting (2.25 Volts); it then
chose a direction towards the center of the lens travel range and proceeded, as described in
Chapter 2, towards shorter focal distances, taking measurements at intervals of .25 Volts
(
Displayed. Once it was determined that the peak had been passed (as a result of a line fitted to
the last five data points changing the sign of its slope and 'tipping over') the algorithm fitted
a parabola to the last twenty points and declared 4.9 Volts (1.13 meters) to be the point of
optimal focus. It should be noted that there is only .3m difference in the position of the focal
plane between a Lens Position of 4 Volts Displayed (1.29m) and 6 Volts Displayed (O.99m)
and little if any noticable change in the video image as viewed on a conventional video
monitor. Also, there is some question as to where exactly the distance inscribed on the lens'
focal ring is measured from: the end of the lens or the position of the image plane on the
-<>
camera tube. In any case the algorithm did better than the author could.
Figures 4-4 through 4-6 show the the Lens Characterization Target and the results of
using the algorithm on it. This is a much more challenging target as there are large amounts
of blank space (and thus more chances for noise to disturb the FMF calculations) and few
edges (only the drawn diagonal lines and circle). Please note the difference in scale on the
FMFaxis between Figures 4-5 and 4-6 and Figures 4-2 and 4-3; the FMF is proportional to
the number of edges present in the image. Despite the minimal information, the algorithm
managed to find a good FMF peak at 3.2 Volts (1.5m). Again, this is better than the author
could do with the video monitor and is very close to the actual distance from the target to the
lens (nominally 1.5m).
For both the Test Grid and the Lens Characterization targets the median filter was set
to a width of N=1 in both dimensions. This setting sucessfully removed the spike noise
caused by the image intensifier (usually one pixel in size as mentioned in Chapter II) without
also removing the lines on the targets, which were approximately 3 pixels wide. Also,
sixteen video frames were averaged to make each image used in FMF calculations.
The third target is shown in Figure 4-7: a plaster cast of a human face (life sized, 2
meters from the camera). This is the most difficult target of the three - almost a 'worst case
test' because there are no definite edges to focus upon and because this target has a three
dimensional surface, whereas the others were flat.
37
Due to the nature of the target and the cameras, noise effected the FMF calculations
far more for this target then even for the Lens Characterization Target. Figure 4-8 is a
composite for 5 seperate trials of measuring the FMF versus the Lens Position for the cast;
note the wide discrepancies in FMF readings near the point of optimal focus (somewhere
between 2 and 3 Volts). Again sixteen frames were averaged for each image used in
calculations, however the median filter was not used. Figure 4-9 contains the same data; the
(
dots are on the mean of the five measurements and the error bars show the extent of the
highest and lowest measurement for that lens position. It must be noted that it was very
difficult to optimally focus on this target manually. Unfortunately the software which applied
the autofocus algorithm so successfully to the previous two targets could not handle the wide
deviations in FMF for this target and perform the hill climbing technique successfully; the
results were extremely erratic. Another test run, for which the data is not shown here, was
performed with a median filter of N=l employed; the results were similiar.
The performance of the algorithm in this case could have been improved
~ with better
.
filtering and extensive averaging of seperate FMF calculations. Unfortunately there was not
sufficient time available to experiment further with the seperable median filters and/or other
methods of reducing the noise and subsequent variations in FMF.
Aside from the noise, two other factors affected the FMF calculations in these trials.
First, the depth of field of the lens (please refer to the section Special Nomenclature), which
depends upon the iris setting of the lens and thus upon the amount of light available to the
camera, was very small. Therefore the front of the cast (the nose) focused at a slightly
different lens setting than the back (the cheeks). Part of the flatness of the left end of the
curve (where there should be a distinct peak) in Figure 4-9 is caused by various portions of
the image sliding into and out of focus at different times in the lens travel; since the FMF is
an average (more or less) of the gradient peaks over the focusing area within the cursor box,
one small area coming into focus and going towards a gradient peak will be canceled by
another which is going out of focus and has gone past its peak. This did not matter with the
previous two targets mentioned because they were flat and the entire target was in focus at
the same time. Second, as mentioned above, the algorithm did not have definite lines and
sharp edges to work with; only relatively smooth slopes in image intensity caused by
shadows and surface texture were available, around the eyes and lips for instance. Smooth
slopes cause wide, low 'hills' in the gradient image, and the peaks of these 'hills' will not be
very far, relatively, above the noise in the image.
38
y
Figure 4-1: Grid Target
70·~--------------------------------------------------------------,
60
400
300
.2001~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
o
2
3
4
S
6
Lens Position (Volts on Display)
Figure 4-2: Lens Position versus Focus Merit Factor Curve for Grid Target
7
39
70'~----------------~--------------------------------~
Parabola:
y
•
= - 734.98
+ 565.07x
-57.61x"2
Preliminary points
Points included in
parabola
••
•
•
•
PO·
(
o
1
.'
234
5
Lens Position (Volts on Display)
Figure 4-3: Results of Autofocus Algorithm for Grid Target
Figure 4-4: Lens Characterization Target
6
7
40
190~----------------------------------------------------~
180-
Ell
Ell I:J
Ell I:J
Ell
I:J
I:J I:J
170(
~
~
~
I:J
Ell
Ell
I:J
I:J
160-
I:J
I:J
1::1
I:J
Ell
I:J Ell
150
Ell
Ell
I:J
.
140
0
I
1
2
3
4
5
6
7
Lens Position (Volts Displayed)
Figure 4-5: Lens Position versus Focus Merit Factor for Lens Characterization Target
200,---------------------------------Parabola
~
fitted to last 20 data points
y = 97.2299+ 50.9316x- 7.9555x1l2 R = 0.96
I!J
Points Used in P arabo la
o
Preliminary
Points
Found Peak
(=3.2V, 1.53m)
1001~~TTTT~~~~~~~~rr~~~~TT~rr~~~~~~rr~~~~~~
o
2
3
Lens Position
4
5
6
(Volts Displayed)
Figure 4-6: Autofocus Algorithm Results for Lens Characterization Target
7
41
(
Figure 4-7: Plaster Face Mold
240
220
200
~
~
~
180
160..:
140
0
1
2
3
4
5
6
Lens Position (Volts Displayed)
Figure 4-8: Lens Position versus Focus Merit Factor for Face (5 Measurements)
7
42
240~----------------------------------------------
220-
~
~
~
:>
-e
__ ~
~
~
!i ii !~
• I~
200
!•
I
I I
180
i • i
I •
I
160..;
140+T~rn~~~rn~~~~Tn,~~~rn~~~~~~,nT~Tn~
o
1
2
3
456
Lens Position (Volts Displayed)
Figure 4-9: Lens Position versus Focus Merit Factor for Face, with Error Bars
7
CHAPTER V
CONCLUSIONS AND RECOMMENDATIONS
Conclusions
On the whole, the algorithm for automatic focusing presented in this thesis works as
anticipated. Within the limits of the work done on pre-filtering of the image, the calculation
of the Focus Merit Factor, its characteristic behavior following a change of the lens' focal
setting, its response to various types of targets - as reported in the previous chapter, and the
method for fmding the peak of the FMF curve all behave as anticipated. The recursive
median filters also worked as expected, although their behavior was not entirely satisfactory
in extreme cases, as with the plaster face cast presented in the previous chapter, and needs to
be augmented. Finally, the method for calculating the threshold above which peaks would be
counted towards the FMF (see Chapter 2) was not very satisfactory at all - it was probably
the major contributor towards the poor results with the face cast.
Recommendations
Keeping in mind the results obtained with the lens characterization target and the
plaster face mold described earlier, there are several areas that need further work. First, a
study needs to be performed to determine the best ways of calculating the threshold.
Currently this threshold is calculated as the average height of peaks due solely due to noise in
the image. As there will be a fair number of peaks greater than this average, the threshold
could be set higher than it is. A statistical method using more sophisticated techniques is
needed. Second, a study needs to be made on optimizing both the gradiant operator and the
median filter for the amount of noise and number of peaks present in the original image.
43
44
When the operator and filter are large they filter out practically all of the noise in the image,
especially shot noise due to the image intensifier, but they also filter out potential edges
which can be used in focusing. When operator and filter are small, they leave the edges, but
also leave noise. Shazeer and Harris (Shazeer and Harris 1985, 151) calculate the shape of
the gradient operator (one dimensional in their case) statistically, based upon the amount of
noise present. The same could be done for the two dimensional operator presented in this
paper. Finally, the number of points that are used in calculating the peak of the FMF curve
can also be adjusted depending upon the relative shape of the FMF curve. For a peak such as
that shown in Figure 4-2, where there are many edges in the original image and the FMF
curve has a relatively sharp peak, a lesser number of points may be used. For curves such as
those shown in Figure 4-10, where the FMF curves are relatively flat and the peak is difficult
to find, a greater number may be used.
APPENDIX
DERN ATION OF TWO-DIMENSIONAL GRADIENT FIL1ERS
The Two-Dimensional Gradient Filter is an operator which, when convolved with an
image, produces a gradient image. Since the image to begin with is a discrete array, the
elements of which sample a continuous two-dimensional function (the 'real' scene), the
gradient image is also a discrete, digital array, the elements of which sample the 'real'
gradient of the continuous scene.
Similar to the familiar first derivative of a one-dimensional function, the gradient is
a measure of the 'slope' of the two-dimensional function (image) at a given point. In
addition, however, the gradient (written as grad!
for a function! (x,y))
showing the magnitude of the slope and the direction. Given a function!
is a vector,
(x,y),J (xO,yO)
increases most rapidly in the direction of grad! (xo,YO).One of the most important
properties of the gradient is that it is defmed as the vector sum of the first derivatives of the
function in the x and y dimensions (Ellis and Gulick 1978, 813):
grad! (xO,yO) = !x(xO,yO)i +!y(xo'YO)j.
(Eq. 1)
Since the gradient is a vector it can be broken down into a magnitude and a direction:
gradMf (xO,yO) = ..J( l!fx(xO,yO)iIl2+ l!fy(xO,yO)jI12)
(Eq. 2)
and
gradif
(xo,YO) = arctan ( l!fx(xO,yO)ill/ l!fy(xO'YO)jll).
(Eq. 3)
In the autofocus algorithm detailed in this report we are concerned primarily with the
magnitude of the gradient at any point; the angle can, however, be easily computed at any
time.
45
46
The derivative of a function!
(x) is defined (Ellis and Gulick 1978, 154) as
! '(xO) =limith"'O(!
(xO+h) -! (xO) )/h.
(Eq.4)
When one is working with sampled functions, such as a digitized image, this can be
approximated with
! '(xO) z (! (xo+M 1(xo) )/f:.
(Eq.5)
where f:. is the seperation between samples (pixels) and! (xO) is the value, or intensity, of
the sample datum (pixel). For images, f:. is customarily set to unity so that the slope in the x
direction.ji, (the x direction is customarily considered to be parallel to the horizontal scan
lines of the video screen/camera), at a pixel (xa'Yb) equals (see line A-Bin Figure A-1)
(Eq.6)
This result is the same as if the scan line
{f (xa-n'Yb);···;j
(xa-1'Yb);! (xa,yb);!
(xa+1'Yb);···;! (xa+n'Yb)}
had been convolved with the operator
{-1,1}.
Please note that all of the results arrived at below for the x gradient apply equally to the y
gradient when the operator is rotated 90 degrees counter-clock-wise. For instance, the x
gradient operator
-1
o
1
-1
o
1
-1
o
1
described below converts to the y gradient operator
111
000
-1
-1
-1.
Since image processing traditionally uses square operators, for simplicity in
programming, the gradient operator is usually expressed as the pair
o
1
1
0
-1
0
0
-1
(also known as the Roberts operator) which, when convolved with the image calculate the
gradient in axes of 45 degrees (I') and 135 degrees ('\) to the horizontal (Ballard and Brown
1982,77).
47
The above method for calculating the slope at a point works well for error free data,
but what if there is an erroneous datum? For instance, if point a in Figure A-I is off of the
'true' value by some error ~e then the slope at (xa'Yb) will be in error by -~e and the slope at
(xa-I'Yb) will be in error by +~e :
·1 x'(xa,yb) =1(xa+I'Yb) -if (xa'Yb)+~]'
(Eq. 7; line A'-B', Figure A-I)
One approach is to defme the slope at (xa;Yb) to be
(Eq.8)
as illustrated in Figure A-I, line C-D; this will result in zero error for the slope calculated for
point (xa'Yb) and errors of + 1/2~ for (xa-I 'Yb) and -1/2~ for (xa+ I 'Yb)' For a function
where the 'true' slope is fairly constant over a length of three pixels or more this method
approximates it well; when an erroneous pixel value is encountered the derivative is
smoothed. The process can be expanded to
(Eq.9)
as needed; however, any peaks or other features in the (non-noisy) data which are less than
N pixels wide will also be smoothed out.
Noise in the pixels surrounding (xa,yb) can be smoothed further by averaging the
slopes of several of the above lines; e.g.
1 x'(xa,yb) = [ if (xa+ I 'Yb ) 1(xa-I,yb ) )/2 + if (xa+2'Yb) -1 (xa-2'Yb) )/4] /2.
(Eq.lO)
This can be re-expressed as
1 x'(xa'Yb) = if (xa+ I,yb ) -1 (xa-I,yb)
) /4 + (1 (xa+2,yb) -1 (xa-2,yb) ) / 8,
(Eq.11)
which is equivalent to convolving the scan line with the operator
{-.5,-I,O,I,.5}.
The errors in this case are ±l/4~ for 1x'(xa±I'Yb)
error of ~ at pixel (xa,yb)'
and ±l/8~ for 1x'(xa±2'Yb) given an
48
Given an image A, expressed as an N by M array of pixels
(xO'YM )
(xa,yM)
(xN,yM )
A=
(xo,yo)
the above can be extrapolated into two dimensions by averaging not only the slopes of lines
through the pixels on either side of (xa,yb) but the slopes of equivalent lines through
neighboring pixels on scan lines above and below (xa,yb)' For instance, one could use the
equation
/ x'(xa,yb) = ( if (xa+1,yb+1) -/ (xa-1'Yb+1)]/2
+
if (xa+ 1'Yb) -f (xa-1,yb) ]/2 +
if (xa+1,yb-1) -/ (xa-1,yb-1)]/2)
(Eq.12)
/3
which is equivalent to convolving the image with the operator
-1/6
0
1/6
-1/6
0
1/6
-1/6
0
1/6.
Since image processors usually work only with integer values, this can be reexpressed as
-1
0
1
-1
0
1
-1
0
1,
also known as the Prewitt operator (Ibid.), if the resulting ji,' (xa,yb) is divided by six. A
famous variation of this is the Sobel operator (Ibid.).
-1
0
1
-2
0
2
-1
0
1
49
which gives additional weight to the slope of the line through the central point.
The 3x3 operator described in equation 12 above can be expanded in various
fashions to allow for greater smoothing and noise reduction. For instance, expanding
equation 11 in the fashion of equation 12 one pixel in each direction results in the equations
f x'(xa'Yb) =
[(xa+ 1,yb+2 ) -f (xa-l,yb+2 ) ) /4 + (f (xa+2,yb+2) -f (xa-2'Yb+2) ) / 8+
if (xa+('Yb+l ) -f (xa-l,yb+l))
/ 4 + (f (xa+2,yb+l) -f (xa-2,yb+l))
/ 8+
if (xa+l'Yb) 1(xa-l'Yb)) / 4 + (f (xa+2,yb) -f (xa-2'Yb)) / 8+
if (xa+ 1,yb-1 ) -f (xa-1,yb-1 ) ) /4+ (f (xa+2,yb-1) -f (xa-2,yb-l ) ) / 8+
if (xa+ 1,yb-2) -f (xa-l,yb-2) ) /4+ (f (xa+2,yb-2) -f (xa-2'Yb-2)) / 8] / 5
(Eq. 13)
which is equivalent to the 5x5 operator
-.025
-.05
0
.05
.025
-.025
-.05
0
.05
.025
-.025
-.05
0
.05
.025
-.025
-.05
0
.05
.025
-.025
-.05
0
.05
.025,
or, reexpressed in integer values through multiplication by 4 as with the Prewitt operator
shown above,
2
1
2
1
-2
o
o
o
2
1
-1
-2
o
2
1
-1
-2
o
2
-1
-2
o
2
1
1.
-1
-2
-1
-2
-1
There is, by the way, no particular reason for the operator to be square except for
convenience. An operator such as
50
-1
0
1
-1
0
1
-1
0
1
-1
0
1
-1
0
1
would smooth out 'wavy' irregular edges in an image while not doing much smoothing to
the gradient normal to the edge, whereas the operator
.
(
-1
-2
0
2
1
-1
-2
0
2
1
-1
-2
0
2
1
would do the converse. In the autofocus algorithm described in this paper the square 5x5
operator described in equation 13 was chosen as a balance between the two and because
video noise in the cameras used in the laboratory tends to cover few pixels while the edges of
interest are many pixels wide.
f(x)
DISCRETE (DIGITAL) SIGNAL
~1Ir"'I~
~e r--+-------
x
Figure A-I: Example of Slope Calculations
SELECTED BIBLIOGRAPHY
Arce, G. R.; Gallagher, N. C.; and Nodesr'T. A. "Median Filters: Theory for One- and
Two-Dimensional Filters," in Advances in Computer Vision and Image Processing,
Vol.2. ed. by Thomas S. Huang. Greenwich, Connecticut: JAI Press, 1986.
Ballard, Dana H. and Brown, Christopher M. Computer Vision. Englewood Cliffs, New
Jersey: Prentice-Hall, Inc, 1982.
Blase, W. Paul. "Laboratory Peripheral Control System for the Numerical Stereo Camera."
in Proceedings of the IEEE 1987 National Aerospace and Electronics Converence
NAECON 1987, by the Institute of Electrical and Electronics Engineers. New York:
Institute of Electrical and Electronics Engineers, 1987, 1528-1540.
Bracewell, Ronald N. The Fourier Transform and its Applications. New York: McGraw-Hill
Book Company, 1978.
Denstman, H. "Die Methoden der Automatischen Scharfeinstellung." ("Methods of
Automatic Focusing") das Elektron, No. 11 (1980),335-337.
"Distance Sensing Uses Automatic Focusing Technology." Sensor Review, Vol. 4, No.4
(October 1984),172-174.
Ellis, Robert and Gulick, Denny. Calculus with Analytic Geometry.New York: Harcourt
Brace Jovanovich, Inc, 1978.
Erasmus, S. J. and Smith, K. C. A. "An Automatic Focusing and Astigmatism Correction
System for the SEM and CTEM." Journal of Microscopy, Vol. 127, Pt. 2 (August
1982), 185-199.
Haase, Hans-Joachim. "Automatische Scharfeinstellung Moderner Videokameras."
Funk-Technik 40, Heft 3 (1985), 107-109.
Hanma, Kentaro; Masuda, Michio; Nabeyama, Hiroaki; and Saito, Yohei. "Novel
Technologies for Automatic Focusing and White Balancing of Solid State Color Video
Camera." IEEE Transactions on Consumer Electronics, Vol. CE-29, No 3 (August
1983), 376-381.
51
52
Jarvis, R. "Range Finding Techniques for Computer Vision." IEEE Transactions on Pattern
Analysis and Machine Intelligence, Vol. PAMI-5, No.2 (March 1983), 127-128.
Kasturi, Rangachar and Walker, John F. "Nonlinear Image Restoration in Signal-Dependant
Noise," in Advances in Computer Vision and Image Processing, Vol.2. ed. by Thomas
S. Huang. Greenwich, Connecticut: JAI Press, 1986.
Krukar, Richard. "Estimating Kernels from Amplitude Spectra." ESD: The Electronics
System Design Magazine, DecembeYl987, 35-36.
Kokoulin, F. 1. and Poleshchuk, A. G. "Automatic Focusing Servo Mechanisms." Soviet
Journal of Optical Technology, Vol. 46(8) (August 1979),466-468.
Matozaki, T. "Image Processing Techniques - Focus on Software." Japanese Journal of
Medical Electronics and Biological Engineering, Vol. 21, No.6 (October 1983),
478-485.
Northrop Corporation. Maintenance Manual for Lens Control Subsystem, Report for
Contract No. F33615-83-C-1096, Project No. FY1175-83-01803, Air Force Wright
Aeronautical Laboratories, Wright-Patterson Air Force Base. Northrop: Anaheim, CA,
1984.
Schlag, John F.; Sanderson, Arthur C.; Neuman, Charles P.; and Wimberly, Francis C.
Implementation of Automatic Focusing Algorithms for a Computer Vision System with
Camera Control, Technical Report CMU-RI-TR-83-14, Carnegie-Mellon University,
August, 1983. (Distributed by the Defense Technical Information Center, Alexandria,
Virginia).
Shazeer, Dov and Harris, Marjorie. "Digital Autofocus Using Scene Content." In SPIE
Proceedings, VoL 534 (1985), 150-158.
Shul'man, M. Ya. "Mechanisms for the Automatic Focusing of Still and Motion-Picture
Cameras." Soviet Journal of Optical Technology, Vol. 45(3) (March 1978), 188-193.
Tee, W. J.; Smith, K. C. A.; and Holburn, D. M. "An Automatic Focusing and Stigmating
System for the SEM." J. Phys. E: Sci. Instrum., Vol. 12 (1979),35-38.
Wolpert, H. D. "AutoranginglAutofocus:
August, 1987, pp 127-130.
A Survey of Systems, Part 2." Photonics Spectra,
0
You can add this document to your study collection(s)
Sign in Available only to authorized usersYou can add this document to your saved list
Sign in Available only to authorized users(For complaints, use another form )