Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Almost all the benefits mentioned in the video are (a) lack of post-processing and (b) high dynamic range. Is that what "log" means in videography?


Log is lower contrast so it's less likely to clip (be a fully saturated color or pure white or black). And clipping inherently limits your max dynamic range.

Log also means a "look" is not baked into the image so, since you're starting from scratch, it's 1) easier to tweak the images so you can cut between two cameras from different manufacturers without distracting differences and 2) you can give the image more of your personality.

As a general note, I've found that in the world of "cinematography", tech terms aren't used very rigorously and there's a lot of cargo cult which comes from the benefit of one tech being conflated as a benefit of something else. It's often hard to sift through the noise when learning.


But clipping occurs before the log transformation. The sensor's ADC is still linear and has a fixed dynamic range regardless of the output encoding format.


Significant clipping happens there yes but more clipping happens when the "look" is applied and contrast is added.


In videography the term "log" is heavily overloaded and you'd want to ask for more detail in order to figure out exactly what is meant.

A pixel value, be it integer or floating point, means little on its own. There's a context for that value which is a color space. In the typical process, you have several color spaces in play: the camera has one for capture. There's one for color processing (the "working" space). And there's one for the display. When a pixel goes through the pipeline, it's processed via color space transformations.

In the "classic" color spaces, the pixel values have a linear relationship, and all of them carry the same amount of information. The "log" color spaces all have a non-linear (gamma) curve: they retain less information at very low and very high pixel values, but subsequently retain more information in the middle. It's a form of compression.

The human eye doesn't respond equally to all levels of brightness, so throwing away detail at the ends for more detail in the middle is usually a great choice. We retain information in the signal at the brightness level where the eye is able to perceive small details and texture, while throwing away information in the signal where it isn't.

We can now map more dynamic range into the same amount of bits, due to our non-linear compression. How large a dynamic range is given by underlying color space we are operating in.

If you go up in camera quality, you will typically see pixels use 10bits or more for their values. Combined with a log-curve, this leads to more information density, which allows capture of an even higher dynamic range. In turn, post-processing can now fix e.g. exposure to a much larger extent.

Finally, a LUT is linear approximation. "Real" color space transformation will use the underlying mathematical curves for much greater precision.


>In the "classic" color spaces, the pixel values have a linear relationship

Almost all color spaces (e.g. srgb or in the video world bt.1886) use a non-linear transfer function? I believe the difference is that gamma and log are different types of nonlinearities, see https://news.ycombinator.com/item?id=37843965 and see https://www.researchgate.net/figure/Linear-response-dashed-v... for an image comparing the two with "linear"

What's confusing is that many people will justify gamma encoding with the claim that visual perception is "logarithmic". I think this is misleading, because the perceptual justification is actually a power-law (Steven's Power Law) as contrasted by the opposing view that perception is logartihmic (Webner-Fechner law, see https://www.appstate.edu/~steelekm/classes/psy3203/Psychophy...). In practice I believe the actual justification for it was that it happened to match the transfer function of CRTs, and these days it's mostly kept around for compatibility, and as an optimization to avoid wasting bits (whether or not it truly fits the human model of perception).

As mentioned by https://computergraphics.stackexchange.com/questions/10315/t..., the real reason why log encoding is nice is because each stop of light gets roughly the same amount of bits. (Log encoding probably also isn't too bad a fit in terms of perception. In an alternate world where we weren't burdened by CRT baggage it'd be a replacement for the now-standard 2.2-esque power gamma).

Also the only reason why log encoded video "looks flat" is because traditional video workflows are not ICC color managed. If you properly applied the inverse transfer function (as any color managed system would automatically do) to display it on an e.g. sRGB screen, the video would appear close to what it did in real-life.


I also found this page which seems to have accurate information (and perhaps the only page I could find acknowledging that compares both log and gamma encodings): https://imatest.atlassian.net/wiki/spaces/KB/pages/114161421...


No, ”log” just means some form of logarithmic response curve when encoding color data. You don’t necessarily get better dynamic range per se, but you get a more useful distribution of the light samples your sensor is taking.


> Is that what "log" means in videography?

> a) Lack of post processing

No. Absence of processing (modifications to make it look 'better') is the default for all non-consumer devices.

> b) high dynamic range

Yes. In practice log is about choosing which bits of color information to retain and which to throw out, to optimize for space.

Log optimizes for retaining detail in very dark and very bright areas by sacrificing detail in the midtones.

Non-log optimizes for midtones. That's all it is.

So if you have a high contrast scene (bright blue sky, someone sitting in the shade), you'll want to use log. In an average/regular contrast scene, you use non-log, that way you get more detail in the midtones.

In photography, there is no need to optimize for space (video is at least 24 frames/sec, photography is a few frames/sec at most, usually), so log is not a thing - we just capture all the things, all of the time.


Thanks for contributing this comment. It's what finally made it click for me, and how logarithms might come into the picture. Especially the difference between video and still photography was helpful to me as a still photographer!

So basically, when you can't afford the space/bandwidth requirements of "raw" data for video, you need to convert the sensor readings to an actual video format right away (the equivalent of "shooting jpeg" on a still camera.)

If you do that conversion using a monotonic concave function (e.g. log) you do get an actual video, but it looks crappy because the tones are not what we would expect. However, it also retains more of the low-end distinctions of the raw data so it's more flexible in processing.

Hypothetically, I could do the same with still photography, by taking the raw data and converting it to a crappy-looking jpeg and distribute that to someone else, who would then have more freedom in processing than with a regular jpeg, but less than the raw data. I think I got it!


A little bit. The log format is non linear. This means there are more details in the shadows relative to the really bright areas. This mimics the human eye and brain which also do not have a linear range of sensitivity.

Basically, the common unit of light in cameras (a stop) is one click on the aperture wheel. E.g. going from 1/11 to 1/16 halves the amount of light. Some cameras of course have a few settings in between. It looks linear to us but it is effectively logarithmic. The dynamic range of the human eye is much larger than the typical camera, screen, or print medium. The human eye has a range of about 20-22 stops (between black and white). A good camera might get between 12 and 14 stops. A decent screen might get to something like 8-10 and print medium is more like 5-7. Taking photos and shooting videos involves a lot of creative choices about what looks natural to us. HDR is basically taking and combining multiple exposures in a way that still looks natural to us on a medium that has less dynamic range than our eyes (-ish, a lot of HDR photography looks a bit unnatural for this reason).

Digital photo processing is about compressing and moving light around to make the most of the much more limited dynamic range of the screen or print medium you are targeting relative to the camera that you used to capture that.

When you do that, most of the interesting information is going to be captured in the darker portions of the image. You typically expose for neutral grey values which is only about 18% of the light. That means half of the darker information (shadows) is in that 18% range of values. And the other half is in the brighter part. Except our eyes are much more perceptive of the darker bits. So, a linear format is not ideal to store that. A log format allocates more bits to the dark half and less to the other 82%. That's a good thing because that allows you to do things like brighten shadows and pull out detail there.

The log format does this by applying a log function to the raw sensor readings. That's why the format looks so flat because all the values end up being relatively close to the 18% mark (neutral). You "undo" this by applying a suitable lut that multiplies the values suitably. You deepen the shadows to near black and brighten the bright stuff to near white. The difference is that you now have full control over this process; can move the white, grey, and black points around. And you can apply color math to the log values before you apply the lut. This is not that different from how you'd process a linear format except now your starting point is better as you are using more bits for the darker parts of the image than for the lighter parts. This gives you more of the captured dynamic range to play with in post processing.

The weakness of the iphone is that while it stores log format, it's not really capable of switching between LUTs on camera or while you are shooting. I'm guessing this just takes too much CPU/battery. So, you have to wait until post processing to see what the end result is going to look like. Some high end cameras have a lot of in camera processing that you can tweak in post processing.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: