Most homeowners find out what their cameras can actually do on the worst possible day. The footage exists, the timestamp is right, and the person in it is a small shape that walks up the drive, does something, and leaves. It was recorded in 4K. Nobody can say who it is.
That isn't a malfunction. It's the camera doing exactly what it was set up to do, which was never identification. The clearest public explanation of why comes from a free UK government document.
Five levels of detail, measured in screen height
In 2009 the Home Office Scientific Development Branch published the CCTV Operational Requirements Manual. It is free, it is 62 pages, and it starts with the question every camera decision should start with: what do you need to see, and why? It then gives five standard levels of detail, each defined by how much of the screen height a standing person takes up.
- Monitor and Control, at least 5%. You can follow the number, direction and speed of people across a wide area, as long as you already know they're there.
- Detect, at least 10%. After an alert, you can search the screen and say with high certainty whether a person is present.
- Observe, 25 to 30%. You can see some characteristic details, such as distinctive clothing, and still see what's happening around them.
- Recognise, at least 50%. Defined this way:
Recognise: When the figure occupies at least 50% of screen height viewers can say with a high degree of certainty whether or not an individual shown is the same as someone they have seen before.
- Identify, at least 100%. The top rung:
Identify: With the figure now occupying at least 100% of the screen height, picture quality and detail should be sufficient to enable the identity of an individual to be established beyond reasonable doubt.
Notice the difference between the top two. Recognise means someone who already knows the person can say "that's him." Identify means the picture itself should be good enough to establish who it is, beyond reasonable doubt. Plenty of home footage that feels useful may be Recognise at best: your neighbour might know the face, but that is not the same as the image establishing it. The manual gives a car park as its example: cameras over the lot set at Detect to catch a break-in, and separate cameras at the exit set at Identify to catch the face. Different jobs, different cameras.
Converting to modern cameras
Those percentages were written for analogue PAL video. The manual converts them for digital systems, and the conversion is where it gets useful. It treats PAL as about 400 effective vertical lines (576 lines adjusted by a standard factor for interlaced video). Identification means roughly 400 lines on the person, head to toe. For a digital camera, the manual's table gives the share of frame height a person needs to fill to reach the same detail:
- 1080p (1,080 lines): Identify at 38%, Recognise at 19%, Detect at 4%.
- 720p (720 lines): Identify at 56%, Recognise at 28%.
- VGA (480 lines): Identify at 84%.
The 2009 table stops at 1080p. Extending the same method to a 4K camera (2,160 lines) is simple division, and this is my arithmetic, not the manual's: 400 ÷ 2,160 is about 19%. So a 4K camera can identify someone who fills about a fifth of the frame height. That's a real improvement over analogue, and it's the honest case for higher resolution.
Now put it on a driveway. The manual's reference person is between 1.64 and 1.76 metres tall. Take 1.7 metres. For that person to fill 19% of a 4K frame, the frame can be no more than about 9 metres (roughly 30 feet) tall at the spot where they stand. For 1080p, at 38%, it's about 4.5 metres (15 feet). A camera set wide enough to show the whole forecourt, the gate and the cars is usually showing far more than that at the gate, which puts a visitor there at Observe or Detect, whatever the box says.
The small print that matters more than the table
The manual attaches three caveats to its conversion table. Each one is a way a system that looks fine on paper can fail.
The weakest link sets the resolution.
The resolution being compared reflects the lowest resolution in the chain, not necessarily the display screen resolution.
A 4K sensor that records a lower-resolution stream, or that you review through a lower-resolution app view or export, is not a 4K system for this purpose. The number that counts is the smallest one between the lens and the file the police receive.
It assumes no heavy compression. The manual warns that recorders almost always store worse video than the live view, and that longer retention usually means harder compression ("Best Storage usually means Worst Image Quality," in its words). Then it says this about buying resolution you can't store:
Therefore there is no point in paying for expensive high resolution cameras if you are not prepared to invest in sufficient storage space and instead make use of heavy compression.
It assumes an average-height person, and even meeting the percentage isn't a guarantee:
Equally, there is no guarantee that individuals will be identifiable just because they occupy >100% of the screen. Other factors, such as lighting and angle of view will also have an influence.
That last one matters most at a front door. A camera mounted high to look down on a doorstep sees the top of a baseball cap. The percentage says Identify. The face says nothing.
What UK police ask for now
The Defence Science and Technology Laboratory now publishes the UK government's CCTV guidance. It lists the 2009 manual as the previous operational requirements manual, and its current two-page sheet for owners whose video may go to the police, UK Police Requirements for CCTV Systems (published August 2022), gives no percentages at all. It does something more practical. Its first two quality rules are to specify what you want to see, and this:
View the recorded video, not the live screen, to assess the system performance.
The supporting notes add that the recorded video and the stills sent to a mobile phone are generally of poorer quality than the live display, and it deals with the most common hope in one line: "It should not be expected that enhancement features, such as digital zoom controls, will provide extra detail." The 2009 manual says the same thing more bluntly: if the camera didn't capture the detail, or compression threw it away, it can't be replaced.
The sheet's test is the one I'd recommend to anyone. Get a volunteer to walk through the door or park a car where it matters, record it under normal operating conditions, and look at the recording. Its standard is one sentence: "If you can't see it then it's not fit for purpose." It also asks that video be exported in its native format without further compression, which is worth checking before you need it, not after.
Why this beat cares
I help design the AI security systems for a veteran-owned (SDVOSB) home-security company run by fellow veterans. I do not own that company and earn nothing from this link. Full policy here.
The lesson I keep relearning in that work is that software can't create pixels the camera didn't deliver. People expect AI to rescue a badly placed camera. It can't. Whether a person is detected, tracked, or later identified by a human investigator, the upper bound is set at installation, by where the camera points and how wide its view is at the spot that matters. I wrote earlier about how face-recognition accuracy numbers are measured under lab conditions. This is the same problem, one step earlier: before any algorithm runs, the frame has to contain enough face.
My military background points the same way. In counter-IED work, a sensor was defined by what it could resolve at the range where the threat was, never by its best-case spec sheet. A camera is a sensor. Ask it the same question.
What to do with this
- Decide the job for each camera. Wide cameras for Detect and Observe are fine and useful. But every entrance a person must pass through (gate, front door, garage door) needs at least one camera set up for Identify at that choke point, at face height or close to it.
- Do the walk test and judge the recording. Have someone walk the route, then export the clip the way you would for the police and look at it on a large screen. Also look at it the way you'd actually see it, on your phone.
- Check the recorded settings, not just the camera model. Resolution and compression for the recorded stream, and whether longer retention quietly lowered quality.
- Test again at night. The manual names lighting as one of the factors that decides whether a face is identifiable, so a view that works at noon needs its own test after dark.
What I could not confirm
The percentages are guidance, not a legal standard. The 2009 manual says its categories are image sizes to aim for rather than minimum standards. The 2022 police sheet says there are no definitive performance criteria for video to be legally admissible, and that admissibility is for the court. Both are UK government documents. I know of no equivalent published US government guidance on this and I have not cited any.
The manual is from 2009 and has been superseded. The government collection page now calls it the previous manual. Its conversion method is simple geometry and still works, but its discussion of equipment reflects 2009 technology, and its table doesn't include resolutions above 1080p. The 4K figure and the driveway distances in this report are my arithmetic using the manual's own method, and they assume the ideal conditions its caveats describe.
Nothing here evaluates or recommends any camera, recorder, app or installer.
The signal
Resolution on the box tells you how many pixels a camera has. It doesn't tell you how many of those pixels will land on a face, survive compression, and reach the investigator. The Home Office unit for identification is simple: a person must fill enough of the recorded frame, roughly 400 lines from head to toe, and even then the angle and the light have to cooperate.
So when a camera is sold to you as "4K," the right follow-up question isn't "how sharp?" It's: at my front gate, in the recording I'd actually export, how much of the frame does a person fill?
Sources
- N Cohen, J Gattuso and K MacLennan-Brown, "CCTV Operational Requirements Manual 2009," Home Office Scientific Development Branch, Publication No. 28/09, first published April 2009, ISBN 978-1-84726-902-7. (PRIMARY. Full 62-page PDF opened and read locally from the UK National Archives copy that GOV.UK links to; text compared with an independent mirror and found identical. Source for: the five observation categories and their screen-height thresholds, with the Recognise and Identify definitions quoted verbatim; the car-park example; the PAL 400-line basis (576 lines times a Kell factor of 0.7); Table 2, including 1080p Identify 38%, Recognise 19%, Detect 4%, 720p Identify 56%, Recognise 28%, and VGA Identify 84%; the three caveats, with the lowest-resolution-in-the-chain sentence quoted verbatim; the reference height of 1.64 to 1.76 metres; the no-guarantee sentence and the high-resolution-cameras sentence, both quoted verbatim; "Best Storage usually means Worst Image Quality"; the statement that most recorders will not store images of the same quality as the live view; the statement that data not captured or discarded by compression cannot be replaced; and the statement that the categories are not minimum standards.)
- Defence Science and Technology Laboratory, "UK Police Requirements for CCTV Systems," two-page guidance, published on GOV.UK 10 August 2022, Crown copyright 2022. (PRIMARY. Both pages opened and read in full. Source for: "View the recorded video, not the live screen, to assess the system performance," quoted verbatim; the note that recorded video and stills sent to a mobile phone are generally of poorer quality; the digital zoom sentence and "If you can't see it then it's not fit for purpose," both quoted verbatim; the volunteer walk-through test; native-format export without further compression; and the statement that there are no definitive performance criteria for video to be legally admissible.)
- Defence Science and Technology Laboratory, "Digital imaging, CCTV and video based evidence," GOV.UK collection, published 16 January 2024, updated 22 January 2024. (PRIMARY. Opened and read. Source for the statement that the 2009 manual is now "the previous CCTV operational requirements manual," hosted on the National Archives website, and for the collection's list of current guidance.)
- Onur Oncer, "Does your camera really recognize faces?," The Signal Report 010, and "The camera that detects but can't record," The Signal Report 042. (Earlier reports in this beat on how face-recognition accuracy is measured and on the gap between what a camera sees and what it keeps.)
Scope note: this report explains published UK government guidance on CCTV image detail. The 4K figure and the distance examples are the author's arithmetic using the 2009 manual's method, labelled as such in the text. It is not a site survey, not legal advice about evidence, and not a substitute for a qualified security designer assessing a specific property. No camera, recorder, app, installer or manufacturer is evaluated or recommended. Disclosure: the author helps design AI security systems for a veteran-owned home-security company, as stated in the body of this report, and does not own that company.
Onur Oncer
U.S. Army combat veteran (Counter-IED / Electronic Warfare), peer-reviewed researcher in microwave spectroscopy, and founder & CEO of Shroombiosis. Consults on laboratory operations, AI, and supplement formulation.