"The report put my hydration in the bottom band, so I came in." We hear this often in clinic. Patients were shown a number and a grade on screen and told what that meant they should do, and yet what the number was being compared against rarely makes it into the explanation. This article is not an argument that skin analysis devices are useless. It is about the frame those scores should be read in.
Three-line summary
- What the machine reads is a physical quantity, and what appears on screen is a score converted from it by rules people wrote. The two are not the same thing.
- Returning the same value on repeat and standing in for the actual state of the skin are different questions — reliability vs validity — and the evidence so far has built up mostly on the first.
- So a skin analysis device earns its place less as a tool for ruling on a single moment than as a tool for repeating the measurement under the same conditions and following the change. A score is not a diagnosis; it is a record of that day's measurement.
What a skin analysis device is actually measuring
What the device does is simpler than it looks. Under fixed lighting it photographs the face, or it brings a probe to the skin and reads a physical quantity: the electrical capacitance of the stratum corneum, the amount of water evaporating from the surface, the wavelength distribution of reflected light, the shadows cast by surface relief. Up to here it is physical measurement, and at this stage the machine works fairly consistently.
What the machine reads and what the screen shows are not the same
One distinction opens up here. What the machine reads is a physical quantity, and what appears on screen is a score converted from that quantity according to rules people wrote. The machine will produce a number one way or another, but whether that number counts as a problem is decided against a threshold someone set in advance. Most of what we look at on a report is the second thing.
- Electrical capacitance of the stratum corneum — converted into a hydration score. The capacitance itself is a physical quantity; the point at which it counts as low is set by a rule.
- Water evaporating from the surface — converted into a barrier-related score. It moves along with the room temperature and humidity at the time of measurement.
- Wavelength distribution of reflected light — converted into a pigment or erythema score. Change the lighting and the input value changes too.
- Shadows cast by surface relief — converted into a pore or wrinkle score. The depth of a shadow shifts with the angle of the light and the position of the face.
Why a score feels so settled
The conversion rule does not show up on the report. The screen keeps just the value after conversion, displayed as a tidy two-digit number or a grade. Hearing "your skin looks a little on the dry side" and reading "hydration, lower grade" land with different weight, and that is natural enough. Numbers look more like facts than judgments.
So when you read a report it helps to step back one stage. Which physical quantity did this score come from, what makes that quantity move, and who set the boundary of the score, against which population? The last two usually do not appear on the screen, and they are often exactly what decides the interpretation.
Some information does not stay on the surface
There is also plenty the device cannot measure. When it started, what you had applied before it started, how it shifts with the season — none of that stays on the surface. The value differs depending on whether you cleansed a moment ago or three hours ago, and that context is not shown on the screen.
So a single report has much the character of a single photograph. It captures that moment's surface precisely, but not the story that led up to the moment. In working out the cause of a skin problem, the second is often the decisive part.
If the same value comes back on repeat, is it accurate?
When a measuring device is evaluated, the literature separates two things. One is reliability — does repeated measurement of the same subject return the same value? The other is validity — is that value actually measuring what it set out to measure? The names sit close together and are easy to mix up, but they are completely different questions.
Put in terms of a thermometer
If a thermometer shows 36.0 degrees every time you use it, its reliability is perfect: it applies the same yardstick over and over. But if it shows 36.0 degrees for someone whose temperature is really 38, it has no validity. The reference itself is off, so it cannot stand in for the actual state. Repeatability and accuracy are separate things.
What construct validity means
Within validity, construct validity is a stiffer standard, because it asks whether a measured value properly represents a concept that cannot be seen directly. We use words like skin texture, elasticity and skin quality every day, but pinning any of them to a single physical quantity is hard. When the thing being measured is itself a bundle of elements, claiming that one measured piece represents the whole calls for separate evidence.
The electrical capacitance of the stratum corneum, for instance, is measured very stably. Whether that value represents the clinical judgment "this person's skin is dry" is a separate question. Capacitance reflects the state of the outermost part of the stratum corneum, while the dryness a patient feels is an experience in which barrier function, inflammation and nerve reactivity are tangled together.
What the validation literature says
The broadest validation work on skin measurement devices is the systematic review Langeveld and colleagues published in Skin Research and Technology in 2022. It pooled the reliability and validity of devices measuring skin color, elasticity and texture, and the conclusion is plain. None of the devices included in the review could be judged valid on construct validity grounds.
That sentence does not mean the devices are junk. Evidence that repeated measurement returns the same value has accumulated across several parameters, and it is closer to saying that the evidence that those values stand in for a clinical skin condition has not yet been met. So on the data available up to now, a skin analysis device belongs less as a tool for settling the state at a given moment and more as a tool for repeating the measurement and tracking the change.
How well it matches varies from parameter to parameter
Validity is not one grade stamped on a device as a whole. Results split quite a lot by parameter. A validation study Cook and colleagues published in the Journal of Dermatological Treatment in 2022 compared the readings of a tablet-based facial analysis software against dermatologists' visual assessment across 14 skin characteristics. Overall agreement was 69 percent, and the spread between parameters was considerable.
| Parameter | Agreement with dermatologist visual assessment | Reading |
|---|---|---|
| Erythema | 83.7 percent | Fairly clearly defined by reflected wavelength |
| Wrinkles | 81.6 percent | Corresponds relatively directly to shadow depth |
| Overall average | 69 percent | The figure across all 14 characteristics |
| Oil | 53.1 percent | Swings widely with touch and elapsed time |
These figures were observed in one particular study and do not carry over unchanged to every device and every person. The direction, though, is worth reading.
What 69 percent actually means
Overall agreement of 69 percent means that in roughly seven parameters out of ten, the device and the dermatologist pointed the same way. Turned around, it also means close to one in three went the other way. What matters here is not which side was right, but that a single report holds parameters carrying different degrees of confidence.
So we do not read a report as one flat list. Parameters that optics catch well, such as erythema, we weight more heavily; parameters that swing with conditions, such as oil and hydration, we judge after checking them against symptoms. Even with the same report, knowing in advance how far to trust each line is what keeps the interpretation steady.
Why color and relief are relatively stable
The wavelength distribution of reflected light and the depth of a shadow have clear physical definitions, and they correspond relatively directly to the redness and the groove depth a person sees with their eyes. In the Langeveld review mentioned above, skin color was also the parameter with the thickest evidence base, at 11 studies, 16 devices and 3,172 participants. For recording and comparing changes in redness, a device has an edge over human memory.
At the other end sits oil
Oil is a parameter that swings widely with touch and elapsed time. The surface 30 minutes after cleansing differs from the surface three hours later, and even on the same surface, the oiliness felt at the fingertip and the reflection read optically carry different information. So on parameters as condition-dependent as oil, the device and the clinical judgment tend not to overlap well. This is why the parameters within one report have to be read with different weights.
Scores from different devices are hard to set side by side
This time the data compares machine against machine. In a device-to-device study the same group published in the Journal of Cosmetic Dermatology in 2022, the rating agreement between two facial analysis systems was 67.7 percent. Roughly one rating in three went the other way.
This result is not an argument that one device is better. It is more accurate to read it as the absence, so far, of agreement on what should be scored and by which standard. Devices may measure different physical quantities in the first place, and even reading the same quantity, the rule that turns it into a score differs.
What that means in practice
- We do not string scores from different devices into a timeline — comparing last time's value from device A with this time's value from device B and reading it as "my skin has got worse" can pull away from what actually happened.
- Tracking assumes the same device — reading the direction of a change requires the yardstick to be fixed. Change the yardstick and you cannot tell a change from a difference in yardstick.
- The difference under the same conditions carries more information than the absolute score — how far today's value sits from a value taken three months ago under the same conditions is clinically more usable than what today's number is.
For the record, we use photography and measurement for clinical records and progress comparison, and we look at reports brought in from elsewhere as reference material too. Where the device differs, though, we do not compare the scores directly; we ask first about the conditions that report was produced under.
What is the score on the screen compared against?
For a score to be assigned there has to be something to compare against. Knowing where that reference came from helps a great deal in reading the number.
A cross-sectional study Seo and colleagues published in the Journal of the European Academy of Dermatology and Venereology in 2022 pooled measurements from 8 devices in 434 Korean participants and proposed an objective classification system. The cut-offs this study produced were the values dividing the measurements into tertiles.
What a tertile means
A tertile says where you sit relative to others within this group. All participants were sorted by the size of the value and split into three blocks, and the values at the boundaries were taken as the cut-offs. Which is why it is hard to treat these as a fixed normal range or a separate boundary of pathology.
So a hydration score in the bottom band means sitting toward the lower end within a reference group, not being in a state that needs treatment. In practice, the same hydration value can be low enough for one person to develop dermatitis while for another it is dry enough that acne settles down and the skin is more comfortable. The number can be the same and still mean something different to that person.
Grades like good and poor have the same structure
Reports often carry grades such as good, average and poor alongside the numbers. The grade is the last stage of the conversion described earlier: the measurement is compared against the cut-off, the range is divided, and the band is given a name. So a poor marking is not a ruling of pathology but closer to a marker that the value fell into the band below the cut-off.
Grades read more firmly than numbers, which is where misunderstanding creeps in. In clinic we see results marked in the lower band on one parameter or another often enough, and that marking on its own does not settle treatment. Within the same band, some people have symptoms and some do not, and checking which is which is, in the end, what the consultation is for.
Cell-level skin age is still at a wait-and-see stage
Services that sample corneocytes, analyze protein markers and calculate a biological skin age have been appearing recently through cosmetics retail channels. It is true that this looks like a step beyond a surface photograph.
To date, however, there is no study addressing the diagnostic validity of this approach. Skin age as a concept does not yet have an agreed definition either, and the output shifts with which markers are combined and how many, while that combination is usually not disclosed. For now we regard it as a stage at which judgment should be held. Talking through the symptoms you have now is more practical than the age number on a report.
Why the same face gives a different result
If the question so far was what splits, it is now time to look at why. The cause is not device performance alone. Measurement conditions often account for a much larger share.
Conditions that move the hydration and oil parameters
Hydration-related readings in the stratum corneum shift with the factors below. None of them is unusual in itself; these are simply the things such a reading moves along with.
- Time elapsed since cleansing — the value read just after cleansing and three hours later differs even on the same skin.
- Room temperature and humidity — the amount of water evaporating from the surface is directly affected by the surrounding environment.
- Whatever was applied just before — after a moisturizer or a sunscreen, the surface itself is already in a different state.
- Settling time before measurement — measuring straight after coming indoors is a different condition from measuring after staying a while.
A report that reads dry on skin with acne underneath can be understood in this light. Measured right after cleansing, the capacitance of the stratum corneum can read low, and surface oil and stratum corneum hydration are different parameters to begin with. Less that the device misread, and closer to a situation where a condition the measurement did not capture changed how the result reads.
Conditions that move the photographic parameters
The same holds for surface photography. A small shift in lighting angle or face position changes the shadows, and when the shadows change, the pore and wrinkle readings move with them. This is why devices are built with systems that fix the photography position and the lighting. Put the other way round, two photographs that did not keep to that system are hard to set against each other.
Worth checking alongside a report you have been given
Checking the conditions before looking at the values makes the interpretation much sturdier. The four below are usually not written on a report, but if you note them and mention them at the consultation, they go straight to work on reading the numbers.
- Roughly what time, and how long after cleansing, it was measured
- Whether you had anything applied that day, and whether it was removed before measuring
- How long you had been indoors before measuring
- Whether it was the same device and the same site as the previous measurement
Which site is measured is a variable too
A reliability study Peperkamp and colleagues published in Skin Research and Technology in 2019 validated skin thickness and elasticity measurement with a particular device in 49 participants. On the elasticity parameter, inter-rater intraclass correlation coefficients were spread widely from 0.23 to 0.80, and test-retest intraclass correlation coefficients ran from 0.25 to 0.84.
The two figures are different indicators. The first is how far a value holds when the person taking it changes; the second is test-retest variability, how far it holds when the same conditions are measured again. What matters is that both ranges were wide, meaning that with the same device, the measurement was dependable at some sites and not at others. This is why a single elasticity figure on a report is hard to use directly as evidence of a treatment's effect.
Why history and visual and tactile examination are still the standard
The first thing we do in the consulting room is not to switch on a machine and start taking photographs, but to look, and to ask. That does not mean we do not use device readings. It is a question of order and weight.
What we ask first
- When it started — days or months changes which causes are worth considering.
- How many minutes after cleansing it starts to feel tight — a sense of it in units of time is more concrete information than any measurement.
- How it changes with the season — we are looking for whether there is a repeating pattern.
- Whether any product has changed recently — this is where the order of a change and its timing shows itself.
Through these questions we add change over time to the picture of this moment. Repeating experiences, seasonal variation and the sequence around a product change cannot be obtained from a measurement at a single point in time. In finding the cause of a skin problem, this information about time is often decisive.
Looking with the eye and feeling with the hand
Visual and tactile examination does the same job. The pattern of flaking, the sense of skin thickness, how quickly the skin springs back when pressed and the oiliness of the surface are information judged together in one pass. They come before being broken down into individual scores, so when findings disagree with each other, the disagreement itself can be the clue.
Skin that shines at the surface while flakes catch at the fingertip is one such example. Surface oil and stratum corneum hydration are not moving in the same direction, and split into two lines of scores it is easy to think one of them must be wrong. In practice both are right, and the combination itself is often the information telling us what is happening in the skin right now.
This is not an argument that using a device is wrong
One point to state plainly, so there is no misreading. This article is not taking issue with consultations that use a skin analysis device. The benefits a device brings in records, explanation and progress comparison are clear, and we use it for exactly that. Being able to look at the same screen as the patient is genuinely useful in clinic.
What differs is whether that screen is the starting point of the judgment or the place where the judgment is confirmed and recorded. The validation evidence accumulated so far supports the second use better than the first. That is the reason we put them in that order, and it is not a problem with the equipment but a matter of where the evidence currently stands.
To put it together: the machine reads this moment's surface against a fixed yardstick, while the history and the visual and tactile examination read the passage of time. The two are less substitutes for each other than steps in an order. So the consultation comes first, and the device reading takes its place confirming and recording that judgment.
So when is the device the right thing to use?
Turn everything above around and the place of a skin analysis device becomes clearer instead. Saying there are uses it does not fit means there are uses it does.
- Tracking treatment progress — measure the same person with the same device under the same conditions, and the direction of the change can be read regardless of how absolutely accurate any one value is. The same yardstick is applied each time.
- Records — comparing the redness of three months ago from memory is hard, but held as an image it becomes possible. It is particularly useful on parameters read stably by optics, such as erythema.
- Explanation — why the moisturizing step needs strengthening now, why this session is better spaced out: all of it lands better looked at together on one screen than told in words alone.
Take the course of treating redness with V-Beam or LDM, or of bringing acne under control with PDT or Gold PTT. Over changes that shift a little at a time across a stretch of time, images taken under the same conditions genuinely help. The change is visible to the eye too, but the interval is long for memory to compare. Even then, what the device settles is not the number of sessions or the choice of procedure but which direction the course is heading in.
The uses that are harder to recommend
Settling the state of the skin from a single measured value, or deciding which procedure to have on the basis of that number. The evidence set out here is hard to read as supporting that use. If a parameter marked in the bottom band on a report leads automatically to a procedure aimed at that parameter, it is worth checking once more what symptoms you actually have right now.
If tracking is the aim, set the conditions first
- On the same device — fixing the yardstick comes first.
- With a similar time since cleansing — matching the timing of the previous measurement makes comparison easier.
- With makeup removed and after settling for a short while — allowing time to adjust to the indoor environment is better for consistency.
- At the same site — change the site and the reliability itself can change.
When tracking is the aim, matching the conditions matters as much as the performance of the device. If the conditions move, even a good device struggles to produce comparable values.
It helps to set the interval as well. Measure too often and the wobble coming from differing conditions looks like change; measure too rarely and there are too few points to read a direction from. The right interval depends on what is being dealt with, so fixing the next measurement point at the first measurement is in practice the simplest approach. Here too, the comparison is not against one previous value but across a line drawn through several points in time.
Worth knowing before a measurement
Skin measurement itself is not an invasive procedure, but methods that bring a probe into contact can leave a passing sting or a pressure mark. If the skin is in a sensitive state or there is acute inflammation, it is better to postpone the measurement. And we would ask you not to stop or change treatment on your own judgment on the basis of a measurement result.
The studies cited here are limited in sample size, measurement sites and study populations, and validation data on facial sites in particular has not yet accumulated sufficiently. Interpretation may change with future research. Skin condition and measured values can appear differently according to individual skin conditions, the measurement environment and the type of device, so results vary between individuals.
Finally
The score on a skin analysis screen is a record of that day's measurement, not a diagnosis. Looking at today's score alone, it does not mean much. Set it next to a value taken three months later under the same conditions and the story changes. From that point the number stops being a verdict and becomes tracking, and tracking genuinely feeds into deciding the direction of treatment.
One more thing: the symptoms you feel yourself come before the score on the screen. How many minutes after cleansing it feels tight, which season it worsens in, what changed after which product — note those and tell us, and they become more useful information than any device. Think of the machine as the tool we use to confirm and record that account.
From Consultation Through Treatment
A board-certified dermatologist examines you directly, identifies which layer the cause sits in, sets the device and the parameters, and then carries out the procedure as the same physician. This is not a structure in which a consultant sells the treatment.
The number of sessions, the intervals, the maintenance point and the cost are agreed together before the first session.
Frequently Asked Questions
- The probe touches the skin during measurement. Is there any irritation?
- Skin measurement itself is not an invasive procedure. That said, methods that bring a probe into contact can leave a passing sting or a pressure mark. If the skin is in a sensitive state or there is acute inflammation, it is better to postpone the measurement. Responses vary between individuals, so tell us before measuring if anything is uncomfortable.
- Is it more accurate to skip the device and go by eye alone?
- This is not a case of one replacing the other. The machine reads this moment's surface against a fixed yardstick, while the history and the visual and tactile examination read the passage of time. Repeating experiences, seasonal variation and the sequence around a product change are hard to obtain from a measurement at a single point, so the consultation comes first and the device reading takes its place confirming and recording that judgement.
- Somewhere else told me I have dry skin, and here you are saying something different. Which is right?
- Rather than one being wrong, there may be context the measurement did not capture. Measured right after cleansing, the capacitance of the stratum corneum can read low, and surface oil and stratum corneum hydration are different parameters to begin with. So we ask first when it started and how many minutes after cleansing it starts to feel tight, and judge that alongside what we can see.
- My elasticity figure is different from last time. Did the treatment not work?
- This is a parameter on which one figure is hard to conclude from. In a study validating skin thickness and elasticity measurement in 49 participants, the elasticity parameter showed inter-rater intraclass correlation coefficients spread from 0.23 to 0.80 and test-retest intraclass correlation coefficients from 0.25 to 0.84. With the same device, in other words, the measurement was dependable at some sites and not at others. We look at progress with symptoms, photographs and clinical findings together rather than one line of figures.
- The photographs say I have more pores than last time. Have they really increased?
- It is worth checking the photography conditions first. In surface photography, a small shift in lighting angle or face position changes the shadows, and when the shadows change, the pore and wrinkle readings move with them. This is exactly why devices are built with systems that fix the photography conditions. It is when images taken under the same conditions are compared that a change can be read at all.
- If erythema matches well, why is oil so far off?
- The easier a parameter is to define physically, the more the device and the human tend to overlap. The wavelength distribution of reflected light and the depth of a shadow have clear definitions, and they correspond fairly directly to the redness and the groove depth a person sees. Oil, by contrast, swings widely with touch and elapsed time, which puts it closer to the things a single photograph struggles to capture.
- I was told my hydration score was in the bottom band. Does that mean I need treatment?
- The bottom band generally means sitting toward the lower end within a reference group. In a study that pooled measurements from 8 devices in 434 Korean participants to propose a classification system, the cut-offs were the values dividing the measurements into tertiles. It is hard to treat those as a fixed normal range or a boundary of pathology. The same hydration value leads to dermatitis in one person and causes no trouble in another, so it has to be read alongside symptoms.
- My skin age came out older than my actual age. Should I be worried?
- Skin age is a concept without an agreed academic definition as yet. The output shifts with which markers are combined and how many, and that combination is usually not disclosed. The approach of sampling corneocytes and calculating a biological skin age from protein markers has no study addressing its diagnostic validity either, so it sits at a stage where judgement is held. Talking through the symptoms you have now is more practical than the number on the report.
- Is there anything I should do before being measured?
- If you can, keep the time since cleansing and the indoor environment close to what they were at the previous measurement. Measuring with makeup removed, after settling for a short while, is better for consistency. Where tracking is the aim, matching the conditions matters as much as the performance of the device.
- Should I just skip skin analysis altogether, then?
- It does not need to go that far. Used to follow change by repeating the measurement on the same device under the same conditions, it is useful enough. But if a single reading is being used to fix your skin type or to settle which procedure you have, it is worth telling us once more what symptoms you actually have right now.
- I was measured at two places on the same day and the results were completely different. Is the machine broken?
- More likely not broken. A study comparing the ratings of two facial analysis systems reported agreement of 67.7 percent. Devices differ in how they measure and in the rules that convert a reading into a score, so values diverging is the more natural outcome. That is why we do not recommend setting scores from different devices side by side to read whether things got better or worse.
- My report came back high on oil, but I always feel dry. Which one is right?
- The two results are most likely measuring different things. Surface oil and a feeling of tightness are separate parameters, and skin that carries oil on the surface while the stratum corneum stays dry is not unusual. The timing matters too: the value shifts with how long it has been since you cleansed. Telling us what you actually feel at the consultation is far more useful for the judgement.