← Shot More Than Once

Shot More Than Once · Part 4 of 6

The Sky from the Dark Frame

An exposure bracket is the same scene shot several times, each frame a stop or two brighter than the last. The dark frame holds the sky and the bright one holds the shadows, and no single frame holds both. Blend is the sixth tool in Sharp 1.2, and its job is to make one picture that does. This article is about how it weighs the frames, why it cannot do that the way the paper does on an iPad, and why the aligner had to be told about brightness before it could line a bracket up at all.

Weighing, not mapping

Blend is exposure fusion, after Mertens, Kautz and Van Reeth (2007). It is not HDR in the radiance-map sense: there is no map of the scene’s true light built from the frames, no tone curve to bring that map back into a picture, and nothing to set. Each frame is given a weight at every pixel, and the frames are added up by those weights. The result is an ordinary picture in the frames’ own colour.

The weight is three things multiplied. Contrast is the size of the Laplacian of the frame’s brightness, its four neighbours against four times itself, which is large on edges and texture and small on flat, clipped or crushed areas. Saturation is how far apart the red, green and blue are, their variance about their own mean. Well-exposedness asks how close each channel is to the middle, and is a bell centred on 0.5:

Swift
let contrast = abs(l + r + u + d - 4 * grey[p])
let mean = (red + green + blue) / 3
let saturation = ((red - mean) * (red - mean) + (green - mean) * (green - mean)
                  + (blue - mean) * (blue - mean)) / 3
func exposed(_ v: Float) -> Float { expf(-(v - 0.5) * (v - 0.5) / 0.08) }
let exposure = exposed(red) * exposed(green) * exposed(blue)
// The small constants keep a flat grey patch, which has no contrast
// and no colour, from weighing nothing in every frame at once.
out[p] = (contrast + 0.01) * (saturation.squareRoot() + 0.01) * exposure + 1e-6

A clipped white sky in the bright frame has no contrast, no colour and sits at the top of the range, so all three factors are near zero there. The same sky in the dark frame has clouds, some blue and a value nearer the middle. The weights are then divided by their total, so that at every pixel the frames’ shares add up to one.

weight = (contrast + 0.01) × (√saturation + 0.01) × well-exposed + 1e-6 exposed = exp(−(v − 0.5)² / 0.08), for each of R, G, B, multiplied share = weight ÷ (sum of every frame’s weight at that pixel)

Where seams come from

Added up by weight directly, one pixel at a time, the frames leave seams. Weights can change from one frame to another over a few pixels, at the edge of a sky or a window, and a hard change in weight is a hard change in tone. The paper’s answer is a Laplacian pyramid: split every frame into bands of detail, coarse to fine, blend each band by a weight smoothed to the same scale, and build the picture back up. Coarse tone then changes over coarse distances and fine detail over fine ones, and the seam has nowhere to sit.

That costs memory Sharp does not have. A full pyramid of a 45MP frame, and an accumulator pyramid beside it to sum the frames into, comes to over 2 GB. The first article measured what the iPad (10th generation) gives an app before a run: about 3.0 GB. Most of that would go on one tool’s bookkeeping.

01 GB2 GB3 GBavailable to the app3.0full pyramid blendover 2Blend, measured0.98
On the iPad (10th generation): what an app is given before a run, the full pyramid blend at 45MP (at least this much), and Blend as built, measured on three 45MP frames.

The pyramid on a small grid

The observation that saves it is that seams live at low resolution. A weight that changes sharply between a sky and a ridge is a coarse thing; so is the tone either side of it. So Blend builds the pyramid only on a small copy of each frame, at most 1,536 pixels on its longest side, and does Mertens’ blend there in full. That gives the blended tone of the whole picture, small.

The full-resolution detail is handled separately. For each frame, its detail is the frame less its own small copy brought back up to full size, bilinearly. That detail is added into the result by the same frame’s weight, read off the small grid, one band of rows at a time, so only one band of one frame is ever in memory beside the result:

Swift
let (lr, lg, lb, wt) = Self.sample(low, weight, x: x, y: y,
                                   width: w, height: h)
guard wt > 1e-5 else { continue }
let sp = (source + x) * 3, dp = (y * w + x) * 3
accum[dp] += wt * (pixels[sp] - lr)
accum[dp + 1] += wt * (pixels[sp + 1] - lg)
accum[dp + 2] += wt * (pixels[sp + 2] - lb)

Then the blended tone from the small grid is brought up and added underneath. The weights only vary on the small grid, so at full resolution they are smooth, and the detail layer has no seams of its own to make. Where every frame agrees, because the weights sum to one, the detail goes in at its own strength. The memory is the result plus a few small pictures, as it is for Clean.

Measured on the iPad (10th generation), three 45MP frames take 6.3 seconds and 0.98 GB. On a Mac, three 1334×2000 frames take half a second.

Lining up a bracket

A bracket shot by hand drifts between frames, so Blend aligns, and left to itself the aligner would align nothing. It scores every candidate warp by the mean difference between the warped frame and the reference, and only accepts one that beats leaving the frame alone by 3%. That margin is what protects it from locking onto the wrong repeat of a pattern.

Before you read on

Two frames two stops apart, one drifted a few pixels. Why would a warp that puts it exactly right fail to beat leaving it alone?

Because nearly all of the difference between the two frames is exposure, and exposure is everywhere. Correcting the drift removes only the small part of the difference that came from movement; the rest stays, and a warp that improves the score by less than 3% is thrown away. Nothing was ever aligned.

So before a frame is registered, it is brought to the reference’s brightness. The mean brightness of each frame’s small copy is measured, and the frame is scaled by the reference’s mean over its own, then clamped back into range:

Swift
let gain = CGFloat(referenceMean / max(mean, 1e-4))
let matched = image.applyingFilter("CIColorMatrix", parameters: [
    "inputRVector": CIVector(x: gain, y: 0, z: 0, w: 0),
    "inputGVector": CIVector(x: 0, y: gain, z: 0, w: 0),
    "inputBVector": CIVector(x: 0, y: 0, z: gain, w: 0),
]).applyingFilter("CIColorClamp")
let alignment = StageTimes.shared.measure("align") {
    aligner.align(moving: matched, reference: reference, extent: extent)
}

The matched copy is only for registering and scoring. What gets warped is the original frame, with its own exposure, since that difference is the whole point of the bracket. The map Blend leaves behind is a source map: at each pixel, which frame weighed most.

The brightness measure works the other way round, too. The focus stack already warned when its frames differed by more than 0.35 stops, because a focus bracket should not; that warning now says that if this is an exposure bracket, Blend is the tool for it. And Blend says the reverse when its frames are all about as bright as each other.

From one photograph

As a check on a real picture, three exposures, −2, 0 and +2 EV, were made with Core Image’s exposure filter from one landscape photograph of heather, rocks and a valley. The blend took the sky and the clouds from the dark frame and the heather from the bright ones, with no seam to be seen. They were made from one photograph, not shot as a bracket, so they share one moment; Blend has not yet been run on an exposure bracket from a real camera.

It does not hunt for ghosts. Anything that moved between exposures, a branch or a person, can show twice, and the settings say so. And it is free: an exposure bracket is three to nine frames, under the free tier’s twelve.

What the tests hold

The tests build a synthetic 320 by 240 scene with a hundredfold range of light, photographed at 0, +2 and +4 stops, gamma-encoded at 1/2.2 and rounded to eight bits, as a camera JPEG is. The first version of the highlight test measured the wrong thing: it counted the edges of the bright frame’s clipped patches, and the edge of a clipped patch is itself a hard edge, detail the bright frame never had. It now measures only inside the patches.

  • Inside the bright frame’s clipped patches, where that frame has no detail at all (under 0.002), the blend holds more than half the middle frame’s detail.
  • In the dark frame’s shadows, below 0.2, the blend has more than one and a half times the dark frame’s detail.
  • Blended on an 80-pixel grid, the picture lands within 0.03 mean difference of the full pyramid and keeps at least 0.8 of its fine detail.
  • A bracket drifted by (5, −3) and (−4, 4) pixels comes back, aligned, to under 0.4 of the unaligned error.
  • The map a blend returns is a source map.

Three frames fit in memory however they are blended. Thirty-five 45MP frames of a vanish did not fit on the iPad’s disk, and the next article is about cutting a long bracket into slabs.