Shot More Than Once · Part 2 of 6
A Wrist That Rolls
Before Sharp lays one frame over another it has to put them on the same grid. Until 1.2 that meant a shift and a uniform scale, and no turn, for a reason that was true of the case it was built for: a bracket on a focus rail does not roll. A burst held in the hand and shot over several seconds does, at the wrist. This article is about teaching the aligner to see a turn, the fifth of every angle it missed the first time, and how much of a picture that gives back.
How a frame is lined up
The aligner is built on Vision’s translational registration, VNTranslationalImageRegistrationRequest, which answers one question exactly: how many pixels one image is shifted from another. Vision’s homographic request, which promises a full perspective warp, was tried first and failed a known-shift test, so it is not used.
A single shift for the whole frame would miss focus breathing, the slight change of magnification as a lens refocuses. So the frame is divided into a two-by-two grid, a 768-pixel window is cut from the middle of each cell, and each window is registered on its own. One more window from the middle of the frame gives a plain global shift. Four tile displacements are enough to fit a uniform scale and an offset, and that is what covered breathing.
Nothing is trusted on faith. Every candidate, including leaving the frame where it is, is scored by the mean absolute difference in luma from the reference frame over a 1024-pixel window in the middle, and a warp has to beat doing nothing by 3% before it is used. That margin is the defence against repetitive subjects, fabric or brick, where a tile can lock onto the wrong repeat and look almost exactly as good as the truth.
let doNothing = score(.identity)
var best = Alignment.identity
var bestScore = doNothing * Float(1 - requiredImprovement)
for candidate in candidates where !candidate.isIdentity {
let candidateScore = score(candidate)
if candidateScore < bestScore {
bestScore = candidateScore
best = candidate
}
}That structure is why rotation could be added without risk to anything that already worked. A turn is one more candidate. If it does not beat the others on the frame in front of it, it is not used.
Fitting a turn
The new candidate is a similarity: a scale, a rotation and a shift, the four numbers that move a rigid picture about in the plane. Each tile contributes its centre and where that centre landed, two equations a tile, so four tiles give eight equations for four unknowns, and least squares picks the similarity that explains them best. Written about the centres’ mean, the answer comes from two sums, a dot product and a cross product of the centred points.
for (c, t) in zip(centres, targets) {
let cx = Double(c.x) - mcx, cy = Double(c.y) - mcy
let tx = Double(t.x) - mtx, ty = Double(t.y) - mty
dot += cx * tx + cy * ty
cross += cx * ty - cy * tx
norm += cx * cx + cy * cy
}
guard norm > 1e-9 else { return .identity }
let a = dot / norm, b = cross / norm
return Alignment(scale: (a * a + b * b).squareRoot(),
translation: CGPoint(x: mtx - (a * mcx - b * mcy),
y: mty - (b * mcx + a * mcy)),
rotation: atan2(b, a))The pair a and b are the scale times the cosine and the sine of the angle, so the scale is their length and the angle is their direction. The fit is only offered within limits: a scale change of no more than 2%, which is already more than a lens breathes, and a turn of no more than 3 degrees, since a wrist rolls a degree or two between shots and anything larger is the fit finding something other than the camera turning. Below a ten-thousandth of a radian it is not offered either, since there is no turn there to speak of.
Four fifths of every angle
The first version did exactly that, and on a burst rolled by a known amount it was wrong in a consistent way. The first frame had been turned so that putting it back meant +1.8 degrees; the fit found 1.39. The last needed −1.5; the fit found −1.20. Every angle came back at about four fifths of itself.
The cause is in the tiles. Each one is registered by translation alone, and a turned window is not a translated one: its edges move one way and the other, and the single shift Vision gives back is a compromise. Over a 768-pixel window and a degree or two that compromise is small, which is why it was thought fair, but four slightly short shifts make a consistently short angle.
The fix does not ask more of the registration. It turns the frame back by what was found, measures the tiles again on the corrected frame, and adds on whatever turn is left. The second measurement is of a much smaller rotation, which translation reads far better, and doing it twice is enough because each round takes most of what remains.
for _ in 0..<2 {
let warped = warp(moving, by: turned, extent: extent)
let (again, moved) = tileDisplacements(moving: warped, reference: reference,
extent: extent)
guard again.count >= 3 else { break }
turned = turned.then(Self.fitSimilarity(from: again, to: moved))
}
if abs(turned.rotation) <= maxRotation { candidates.append(turned) }After the two rounds the same frames come out at 1.79 degrees for 1.8, and −1.49 for −1.5.
An alignment you can add up
Refining needs one alignment followed by another to be expressible as a single one, and that is true of similarities: a turn and a scale and a shift, followed by another, is still a turn and a scale and a shift. The alignment is kept as plain data, and composed by going through its transform. In Core Image’s space, with y upward, the transform’s a and d are the scale times the cosine of the angle, b is the scale times its sine, and c is the negative of b.
public var transform: CGAffineTransform {
let a = scale * cos(rotation), b = scale * sin(rotation)
return CGAffineTransform(a: CGFloat(a), b: CGFloat(b), c: CGFloat(-b), d: CGFloat(a),
tx: translation.x, ty: translation.y)
}
public func then(_ next: Alignment) -> Alignment {
let t = transform.concatenating(next.transform)
return Alignment(scale: Double((t.a * t.a + t.b * t.b).squareRoot()),
translation: CGPoint(x: t.tx, y: t.ty),
rotation: Double(atan2(t.b, t.a)))
}The alignment also became Codable. Both of those are needed beyond this article: a saved project has to remember how each frame was moved, and a long bracket stacked in slabs puts a frame on its slab’s grid and the slab on the final one, which is composition again.
One smaller change came with the turn. The aligner reports how far an alignment moves a frame, and that used to be a question about the shift. A rotation about a point moves the centre hardly at all and the corners most, so the measurement now applies the transform to the frame’s four corners and reports the largest movement.
A burst with a known answer
Sharp’s command-line tool has a burst command that takes a real photograph, builds from it the burst each tool exists to fix, and measures how much of the photograph comes back. Because the photograph is the right answer, a change that quietly makes things worse shows up as a number. Rotation got its own option, --roll, which turns each frame about its centre by a set amount more than the last.
let frames = (0..<frameCount).map { i in
let step = i - frameCount / 2
return noisy(rolled(shifted(truth, by: drift * step), degrees: roll * Double(step)),
seed: i)
}The run was sharp burst data/landscape/stack01.jpg --roll 0.3: twelve frames of a 1667 by 2500 photograph, each rolled 0.3 degrees more than the last, which puts the frames at the two ends 1.8 and 1.5 degrees from the middle one, with 4% noise added to each. Clean, the averaging tool, stacked them with alignment on, and the result’s distance from the original photograph was compared with the best single frame’s.
Before you read on
With the aligner fitting only a shift, is the average of twelve rolled frames closer to the photograph than one frame, or further from it?
Further. A roll about the centre barely moves the middle of the frame and moves the corners most, so no single shift puts twelve differently turned frames back, and they are averaged still turned. Averaging frames that do not line up is blur. The result was 0.6 times as close as one frame: worse than not stacking at all.
With the first rotation fit, the four-fifths one, the result was 1.2 times closer than one frame: better than not stacking, but only just. With the refinement it was 2.4 times. The same burst built with no roll at all, which is the ceiling for this test, came out at 3.3 times, against the 3.5 that averaging twelve frames of independent noise predicts.
The gap that remains between 2.4 and 3.3 is not a wrong angle. A turned frame has to be resampled to be turned back, which softens it slightly, and turning a frame brings in edges that were never photographed, filled by smearing the nearest pixels outward. Both cost the average something, and neither is a matter of fitting.
Nothing else moved
A new candidate in the aligner is only safe if frames that do not need it are left as they were. On a burst with no roll and a shift of 3 pixels a frame, --drift 3, the three tools the command measures gave the same numbers after the change as before: 0.0066 for clean, 0.0089 for vanish and 0.0342 for trails.
Adding --roll turned up an older fault. The command-line tool keeps a list of the options that take a value, so it knows not to read the value as a file to stack. --drift had never been on it, so --drift 3 had been taking “3” for a file name. It is on the list now, beside --roll.
What has not been tried is a real rolled burst. Every number above is on a computed roll applied to a photograph. A burst shot over several seconds on a camera, where the wrist actually rolls and the frames bring their own sensor noise and their own small changes, is still untried.
What the tests hold
- A scene of smooth random blobs, 640 by 480, chosen because it has nothing periodic for a tile to lock onto the wrong repeat of, turned by 1.5 degrees, is found to need −1.5 degrees within a quarter of a degree, and once warped back its difference from the original falls below 35% of what it was.
- An alignment composed with another moves a point exactly where applying one and then the other would.
- An alignment with a scale, a shift and a rotation encoded to JSON and decoded again is the same alignment.
- The earlier test still passes: a five-frame focus bracket drifted by a known shift stacks, aligned, to less than 60% of the unaligned stack’s difference from the same frames stacked without the drift.
Rotation is a setting the aligner can find for itself. The next article is about one it cannot: the size of detail a focus stack looks at, and why the answer is four widths, chosen by eye.