The labour theory of value has been tested empirically, repeatedly, across many national economies, using data that governments publish and methods that anyone with the tables and a linear algebra package can replicate. That fact alone puts this article at odds with most writing on the subject, which treats the theory as either a philosophical commitment or a discredited relic and in both cases as something that evidence cannot touch.
The tests exist, the results are striking, and the results are also contested on grounds that have nothing to do with any of the classical objections. The dispute is statistical: because both the price figures and the labour-value figures for each industry are scaled by how big that industry is, a large part of the observed agreement between them may be measuring nothing more than the fact that big industries are big. That objection, and the replies to it, is the whole of the serious argument, and it is essentially invisible outside the journal literature.

What follows states the claim in a form that could be measured, describes exactly how labour values are computed from published national accounts, reports what the studies found and across which economies, states the aggregation critique at its full strength including the demonstration that arbitrary non-labour inputs can produce comparably impressive results, gives the replies, and reaches a split verdict. The strong version of the empirical claim is not established. A weaker version has some support and remains disputed. Anyone who tells you otherwise, in either direction, has not read the exchange.
The claim in a form that could be measured
Before anything can be tested, the proposition has to be stated so that a number could disagree with it. The labour theory of value in its full form does several things at once, and only one of them is a candidate for this kind of measurement. The theory’s other claims, about the origin of surplus and about the social form that labour allocation takes, are not tested by any of the work described here, and readers who want those distinctions drawn properly should start with the labour theory of value explained and with the criticisms of the labour theory of value, which owns the survey of the theoretical objections.
The measurable proposition is this. If the labour time required, directly and indirectly, to produce a unit of each commodity is the magnitude that regulates exchange ratios, then across the industries of a real economy the relative prices of output should stand in a close and systematic relation to the relative labour requirements of that output. Not identity, since the theory itself denies identity: prices of production diverge from values systematically according to the ratio of constant to variable capital in each branch. But if the deviations are as large as the divergence between prices and any arbitrary alternative magnitude, the theory has no empirical content at the level of relative prices.
That is a genuine test in the sense that it could fail. The theory predicts a close relation; the deviation could turn out to be enormous; the labour magnitude could turn out to perform no better than the electricity magnitude or the steel magnitude or a randomly generated vector. The three possible outcomes are analytically distinct and the literature contains claims for all three.
What would count as evidence for the labour theory of value?
Labour requirements tracking relative prices substantially more closely than alternative input measures do, after the influence of industry size has been removed from the comparison. Raw correlations that have not been normalised for sector size do not qualify, because both quantities are scaled by output and will correlate for that reason alone.
How labour values are actually computed
The method is worth setting out properly, because almost every popular treatment of this literature reports a correlation without saying what was correlated, and the whole dispute turns on that.
Every industrialised country’s statistical agency publishes input-output tables, which record what each industry buys from every other industry and what each industry pays in wages, over a given period. From these tables it is possible to compute, for each industry, the total labour required to produce a unit of its output, counting not only the labour employed in that industry but the labour embodied in all the inputs it purchased, and the labour embodied in the inputs to those inputs, and so on through the entire chain. The quantity is called the vertically integrated labour coefficient, and it is obtained by inverting the matrix of technical coefficients and multiplying by the vector of direct labour inputs. The mathematics is standard and identical to that used in mainstream input-output analysis for entirely different purposes.
That gives one vector: labour required per unit of output by industry. The other vector comes from the same tables and is the money value of gross output by industry, which is the price side. The test compares the two.
Several decisions have to be made along the way, and each of them is a place where results can move. Labour has to be measured somehow: as hours worked, as persons employed, or as the wage bill, and the three are not equivalent, since using the wage bill imports the skilled-labour reduction through the back door by treating relative wages as the conversion between kinds of labour. Fixed capital has to be handled, either by treating depreciation as a flow of embodied labour or by more elaborate methods, and the treatment affects capital-intensive industries most. Sectors that Marxian accounting classifies as unproductive, principally parts of finance, trade, and government, have to be either included with the rest or separated out, and the choice matters because these sectors have unusual ratios of labour to output. Imports have to be handled, since goods entering from abroad embody foreign labour under different technical conditions. And the industries themselves are aggregates defined by a statistical classification, so the level of aggregation is a choice: an economy divided into thirty sectors and the same economy divided into four hundred will not give the same answer.
None of this is a scandal. Every empirical exercise in economics involves choices of this kind, and the mainstream measurement of the capital stock involves worse ones. It does mean that a reported result is a result about a conjunction of the theory and a measurement scheme, which is why the sensible way to read this literature is to look for results that survive across schemes rather than to take a single headline number seriously.
What has been measured, and where
The research programme is roughly a generation old at the point where it becomes substantial, and it has been carried out by a small number of researchers working across a number of national datasets.
Anwar Shaikh’s work from the nineteen eighties onward established the basic approach in the anglophone literature and made the case that the deviations between market prices, prices of production, and labour values in real economies were far smaller than the theoretical literature on the transformation problem had led people to expect. Shaikh’s argument was that a debate conducted entirely on the possibility of large deviations had never checked whether large deviations actually occur.
Edward Ochoa’s study of the United States, developed from doctoral work and published in the late nineteen eighties, computed labour values and prices of production for the American economy across several benchmark years and reported close agreement between the two and between both and market prices. This is the study most frequently cited as establishing the result for a large modern economy.
Paul Cockshott and Allin Cottrell carried the programme forward in the nineteen nineties, working with United Kingdom data, and made a contribution that turned out to be more important than the correlations themselves: they compared labour as a value basis against alternative bases, computing what the price structure would look like if some other input, such as energy or a particular material, were treated as the value-forming substance. Their claim was that labour outperformed the alternatives, which is precisely the comparison that any serious test requires and that earlier work had not systematically made.
Replications and extensions have been conducted across a range of economies, including work on Yugoslavia, Greece, Mexico, Sweden, Italy, Japan, and several others, generally by researchers working in the same programme and generally reporting results of the same order. Cross-country studies have attempted to establish whether the pattern is robust to very different industrial structures and levels of development.
The headline finding, reported consistently across this body of work, is that the correlation between labour values and money prices of gross output across industries is very high, commonly reported at levels that the authors describe as strong and that a reader accustomed to social-science data would find remarkable. Precise coefficients vary by country, by year, by aggregation level, and by measurement scheme, and anyone citing a specific figure should cite the specific study rather than the literature as a whole.
Where this research programme came from
The empirical study of labour values is younger than the theory by about a century, and the reason is technical rather than ideological: the calculation was not feasible before the data and the mathematics existed.
The mathematics arrived with input-output analysis, developed by Wassily Leontief from the nineteen thirties onward, which represented an economy as a matrix recording the flows of goods between industries and made it possible to compute the total requirements, direct and indirect, for producing any bundle of final output. The technique was built for planning and forecasting and has been used across the whole spectrum of economics, from wartime resource allocation to environmental footprint accounting. It is not a Marxian tool and its association with this literature is incidental: it happens to compute exactly the magnitude the labour theory of value needs, which is the total labour embodied in a unit of output through the entire chain of production.
The data arrived with the postwar expansion of national accounting. Statistical agencies began publishing input-output tables at regular intervals with increasing sectoral detail, and by the nineteen seventies most industrialised countries had comparable series covering several decades.
The motivation arrived from an unexpected direction. For most of the twentieth century, argument about the labour theory of value was conducted entirely on theoretical grounds, and the transformation problem literature had established that deviations between values and prices of production were possible and could in principle be large. What nobody had checked was whether they were large in actual economies. The theoretical possibility of a large deviation and the empirical fact of one are different things, and a literature that had spent seventy years on the first without examining the second was making an obvious omission. Anwar Shaikh’s central methodological argument, made from the early nineteen eighties, was exactly this: that the question of magnitude is empirical and had gone unasked.
The timing also owed something to the Sraffian challenge. If value magnitudes were redundant for determining prices, as the redundancy argument held, then one obvious response was to show that they nonetheless track prices closely, which would make redundancy a formal result of limited practical importance. Whether that response succeeds is contested, since a redundant magnitude that happens to correlate with prices is still redundant, but it explains why the empirical programme accelerated when it did.
What the programme inherited from its origins is a set of assumptions that came in with the tools. Input-output analysis assumes fixed technical coefficients, constant returns to scale within the period, and a well-defined sectoral classification, none of which is exactly true and all of which are standard. Anyone assessing the results should hold in mind that these assumptions are shared with a great deal of orthodox applied economics and that criticising the Marxian application of them while accepting the orthodox one is inconsistent.
Reading one of these studies: a walkthrough
The most useful thing a reader can acquire here is the ability to interrogate a paper rather than to remember a conclusion. A typical study in this literature proceeds through six stages, and there is a question worth asking at each.
The first stage is data selection: a country, a year or set of benchmark years, and a level of sectoral detail. The question to ask is how many sectors. A study working with thirty broadly defined sectors and a study working with four hundred are not doing comparable exercises, because aggregation averages heterogeneity away and produces smoother relations. Results at coarse aggregation should be read as upper bounds on the agreement between prices and values.
The second stage is computing the labour vector. The question to ask is what labour was measured. Hours worked treat all labour as homogeneous, which the theory denies but which keeps the value calculation independent of prices. The wage bill weights labour by its price, which handles heterogeneity and destroys the independence, since price information is now inside the supposedly independent variable. Persons employed ignore variation in hours. A study that does not state its choice explicitly cannot be assessed at all.
The third stage is computing prices of production, which many but not all studies do as an intermediate step. This requires a uniform profit rate and a treatment of fixed capital, and the question to ask is how depreciation and capital stock were handled, because this bears most heavily on the capital-intensive sectors where the theory predicts the largest divergence between value and price of production. A crude treatment here softens exactly the test the study exists to perform.
The fourth stage is the comparison itself, and this is where the aggregation problem lives. The question to ask is whether the compared quantities are totals or per-unit magnitudes. If the study correlates the aggregate labour content of each sector’s output with the aggregate money value of that output, both are scaled by sector size and the result is uninformative. If the study computes per-unit values and prices and then measures deviation, the objection in its simple form does not apply.
The fifth stage is the choice of statistic. The question to ask is which one, and the answer matters more than most readers expect. A correlation coefficient on levels, a correlation on logarithms, a coefficient of determination from a regression, a mean absolute deviation, a mean absolute weighted deviation, and an angular distance measure between the two vectors will not tell the same story about the same data. Several authors have argued that correlation and regression statistics on this kind of data are systematically flattering, and proposed distance-based alternatives precisely to avoid that.
The sixth stage is the comparison against alternatives, and most studies skip it. The question to ask is whether any candidate other than labour was tested under identical treatment. Without that, a high figure is uninterpretable, because there is no baseline for what a high figure means. With it, the study becomes informative regardless of which way the result falls.
A reader who runs those six questions over any paper in this literature will place it accurately within about ten minutes, which is a better return than any summary of the field can offer.
The wage-profit curve result
There is a second empirical finding in this programme that receives a fraction of the attention given to the price-value correlations and is arguably more consequential. It concerns the shape of the wage-profit curve.
In a linear production model, for any given technique, there is a trade-off between the real wage and the rate of profit: as one rises the other falls, and the relation between them can be plotted. Theory places no restriction on the shape of this curve. It can be convex, concave, or wavy, and the possibility of wavy curves is what generates the phenomena of reswitching and capital reversal, in which a technique that is most profitable at low profit rates becomes most profitable again at high ones, with another technique preferred in between. Those possibilities were the analytical core of the Cambridge capital controversies, and they were used to show that no consistent relation between the capital intensity of a technique and its profitability can be assumed.
The empirical finding, reported for several economies, is that the wage-profit curves computed from actual input-output data are close to linear. Not exactly linear, but far closer to straight lines than theory requires them to be. Ochoa’s work on the United States is the study most often cited for this, and similar results have been reported elsewhere.
Two consequences follow and they pull in different directions. The first is that near-linear wage-profit curves imply that prices of production stay close to labour values as distribution changes, which explains the price-value correlations structurally rather than merely reporting them. If the curve is a straight line, the deviation of prices from values is small across the whole range of possible distributions, which is a much stronger result than a correlation at one distribution. The second consequence is that the theoretical possibilities the capital controversies established, reswitching and capital reversal, appear to be empirically rare. That is uncomfortable for the side that used those possibilities as a critique of marginal productivity theory, since it suggests the critique was logically valid and practically unimportant.
Anyone arguing about this literature should know that the wage-profit curve result exists, because it is the strongest thing in it and because it shifts the argument from statistics to structure. The aggregation critique applies to correlations between size-scaled vectors; it does not apply in the same way to the observed shape of a curve computed from technical coefficients. The critical response has been to question the empirical construction of the curves and the treatment of fixed capital within them, which is a narrower and more technical dispute than the one about correlations.
The aggregation critique
Then comes the objection, and it is devastating in a way that requires no economics at all to understand.
Both vectors being correlated are measured in totals rather than per unit. The labour value of an industry’s output, as computed, is labour per unit multiplied by the number of units. The price of an industry’s output is price per unit multiplied by the number of units. Both are therefore scaled by industry size, and industry size varies across sectors by orders of magnitude: a national economy contains sectors with output measured in billions and sectors with output measured in millions.
When two variables are both multiplied by a third variable that ranges over orders of magnitude, they will correlate strongly whatever the relation between the underlying per-unit quantities. This is a textbook spurious correlation and it does not require any relation whatever between labour content and price. A large industry has a large labour bill and a large output value because it is large.
The critique was pressed most forcefully by Andrew Kliman in work published in the early two thousands, using United States data, and its force comes from a demonstration rather than an assertion: if you generate a vector of random numbers, scale it by industry size in the same way, and correlate it with the price vector, you obtain a correlation of the same order as the one the labour vector produces. Similar exercises have been run using various non-labour inputs, and they too produce impressive-looking results. If a random vector and an arbitrary input do as well as labour, then the reported correlation is measuring the size distribution of industries and nothing else.
Related methodological objections were raised independently by other authors, including work by Ian Steedman and Judith Tomkins in the late nineteen nineties on how deviations between prices and values should be measured at all, which pointed out that the choice of measure and normalisation has a large effect on how big the deviations appear. Nitzan and Bichler pressed a version of the same objection from a different theoretical direction, arguing that the correlations reported in this literature are artefacts of the way the data are constructed rather than findings about the economy.
Why do critics say the price-value correlations are spurious?
Because both the price figure and the labour-value figure for an industry are totals, scaled by that industry’s size, and industry sizes differ by orders of magnitude. Two quantities scaled by the same large-ranging third quantity will correlate strongly regardless of any relation between the underlying per-unit magnitudes.
What the demonstration does and does not establish
The arbitrary-input demonstration is the single most important result in this literature and it needs stating precisely, because both camps have overclaimed from it.
What it establishes is that the raw cross-sectoral correlation between aggregate labour values and aggregate prices, reported without normalisation, is not evidence for the labour theory of value. That conclusion is secure. Any study reporting such a correlation as a finding has reported a property of the size distribution of industries.
What it does not establish is that there is no relation between labour requirements and prices. Showing that a test is uninformative is not showing that the hypothesis it purported to test is false. A defender who responds to the demonstration by producing a normalised comparison has not evaded the objection; they have done what the objection required. And the critic who treats the demonstration as settling the substantive question has committed the mirror error of the defender who treated the raw correlation as settling it.
There is a further subtlety worth registering because it is where the argument becomes genuinely difficult. The size-scaling problem does not affect all measures equally. A correlation coefficient across industries measured in totals is badly affected. A measure of the average absolute percentage deviation between the price and the value of a unit of each industry’s output is not affected in the same way, because it is computed per unit. Studies reporting the latter kind of measure are not vulnerable to the aggregation critique in its simple form, and several of them report deviations that are modest, though what counts as modest is itself disputed and depends on the normalisation chosen. A reader assessing any particular study should therefore ask first which kind of measure it reports.
The replies, and how good they are
The defenders of the programme have made four responses, of distinctly different quality.
The first response is to compute correlations on per-unit or deflated measures rather than on totals, removing the scaling by industry size directly. This is the correct response and it is the one that makes the literature worth continuing. Results computed this way are weaker than the raw correlations, which is expected, and the question becomes whether what remains is substantial. Reports differ, and the difference between studies is largely a difference of normalisation choice.
The second response is the comparative one, and it is the most persuasive available. The relevant question is not whether labour requirements correlate with prices in absolute terms but whether they do so better than alternatives, and the comparison must be run with the same scaling applied to every candidate. Cockshott and Cottrell’s framing of the problem in these terms is the methodological contribution that survives the whole dispute, whatever one concludes about their results. If, under identical treatment, labour outperforms energy, steel, and randomly generated vectors, that is a finding; if it does not, that is also a finding, and a decisive one.
The third response is the replication argument: that the pattern appears across economies with very different industrial structures, sizes, and levels of development, and that a purely statistical artefact should not be expected to produce results of similar magnitude across such variety. This response has some force and it is weaker than it looks, because the size distribution of industries is itself broadly similar across market economies, so an artefact driven by that distribution would replicate for exactly the same reason.
The fourth response is theoretical rather than statistical, and it points to work explaining why near-proportionality between prices and labour values should be expected on structural grounds. Research on the mathematical properties of input-output matrices, particularly the distribution of their eigenvalues, has shown that in real economies the higher-order effects that would drive prices away from labour values are small, because the matrices are strongly dominated by their first eigenvalue. This line of work, developed by researchers including Theodore Mariolis and Lefteris Tsoulfidis, offers an explanation of the empirical pattern rather than merely reporting it. It is the most interesting development in the area and it cuts both ways: if near-proportionality follows from a general structural property of production matrices, then it may be a property of any economy with those characteristics rather than confirmation of a specifically Marxian proposition.
The evidence table
The table below is this article’s findable artifact. It summarises the state of the literature by study family rather than by individual paper, and the last column is the one that matters.
| Study family | Data used | What was computed | Reported result | Principal objection | Status of the objection |
|---|---|---|---|---|---|
| Foundational work on deviations, from the nineteen eighties onward | United States input-output tables, multiple benchmark years | Labour values, prices of production, and market prices compared across industries | Deviations far smaller than the transformation-problem literature had assumed | Deviations measured on aggregate rather than per-unit magnitudes in early formulations | Partially answered by later per-unit measures; original framing conceded |
| National replications across many economies | Input-output tables for Yugoslavia, Greece, Mexico, Sweden, Italy, Japan, the United Kingdom and others | Cross-sectoral correlation between labour values and money output | High correlations reported consistently across very different economies | Cross-country replication does not defeat an artefact if the artefact has the same source everywhere | Open; the reply and the objection are both reasonable |
| Alternative value-basis comparisons | United Kingdom and other national tables | Labour compared against energy, materials, and other candidate bases under identical treatment | Labour reported to outperform the alternatives | Comparison must apply identical scaling to every candidate, which early versions did not always do | The methodologically decisive test; results contested rather than settled |
| The aggregation critique | United States tables, plus randomly generated vectors | Correlation of price with size-scaled random and arbitrary vectors | Random and arbitrary vectors produce correlations of the same order as labour | Showing a test is uninformative does not show the hypothesis false | Secure against raw correlations; does not by itself settle the substantive question |
| Per-unit deviation measures | Various national datasets | Mean absolute or mean percentage deviation of unit price from unit value | Deviations reported as modest, though what counts as modest depends on normalisation | Choice of normalisation and deflator materially affects the reported magnitude | Open; this is where the serious argument now sits |
| Structural explanations from matrix properties | Input-output matrices from several economies | Eigenvalue distributions of vertically integrated coefficient matrices | Near-linearity of price-value relations follows from matrix structure | If it follows from general structural properties, it is not specifically Marxian confirmation | Open, and the most analytically interesting line available |
The aggregation test
The rule this article advances is the aggregation test: no price-value correlation counts as evidence unless the study reports what happens when industry size is normalised out, and applying that single test reorders the entire empirical literature.
The test is deliberately simple, because a rule that requires expertise to apply will not be applied. Confronted with any claim about empirical support for the labour theory of value, ask one question: were the quantities compared per unit of output, or were they totals? If they were totals, and no normalisation was performed, the reported result is uninformative regardless of how high the coefficient is and regardless of how many countries it replicates across.
Applying the test has consequences the defenders of the programme should welcome, because it clears away the weakest evidence and leaves the strongest visible. The studies that survive the test are the ones that compared labour against alternative bases under identical treatment, and the ones that reported per-unit deviation measures rather than aggregate correlations. Those are a minority of what is cited in popular argument and they are the only part of the literature worth arguing about.
The test also disciplines critics. A critic who cites the arbitrary-input demonstration and concludes that the theory has no empirical support has skipped the same step in the opposite direction, since the demonstration bears on the raw correlations and not on the normalised comparisons. The honest critical position is that the strongest evidence offered for the theory is far weaker than its advocates present, not that the question has been closed.
What these studies do not test
Confusion about the scope of this literature is widespread and it inflates the significance of the results in both directions.
These studies do not test the transformation problem. The transformation dispute concerns whether Marx’s procedure for converting values into prices of production is internally consistent and whether the two aggregate identities can be preserved simultaneously. That is a question about the logical structure of a derivation, and no correlation between empirical vectors bears on it. A study finding high agreement between values and prices does not resolve the transformation problem, and a study finding low agreement does not create it. Readers who want that dispute should go to the transformation problem explained, which owns it in this series.
These studies do not test the theory of surplus value. The proposition that profit originates in unpaid labour is not a claim about the correlation between price and value vectors across industries, and an economy in which labour values tracked prices perfectly would be perfectly compatible with any account of where profit comes from. The empirical work relevant to that proposition concerns the measurement of the rate of exploitation and belongs elsewhere in this series.
These studies do not test the social-form claim, which is not the kind of proposition that input-output data addresses at all.
And these studies do not test the marginalist framework against the Marxian one, despite frequent claims to the contrary. A finding that labour requirements track prices is consistent with a marginalist account in which long-run competitive prices equal marginal costs and labour is the dominant component of cost, which is roughly what one would expect in an economy where wages are a large share of value added. Establishing that one framework fits the data does not establish that the other does not, and the empirical result that would discriminate between them has not been constructed. The comparison itself, and the reasons it is so hard to settle empirically, is handled in Marx versus the marginalist theory of value.
The statistics a reader needs, and why they matter here
This dispute is unusual in that a small amount of statistical understanding settles most of it, and the absence of that understanding is why the popular versions are so bad in both directions.
The first point concerns what a correlation coefficient measures. It measures how closely two variables move together in a linear relation, and it is scale-sensitive in a way that matters enormously here. If two variables are each the product of a per-unit magnitude and a size magnitude, and the size magnitude varies over several orders of magnitude while the per-unit magnitudes vary within a much narrower range, the correlation between the products will be dominated almost entirely by the size term. Nothing about the per-unit relation is being measured. This is not a subtle statistical point and it is not disputed by anyone in the literature; what is disputed is how much of the reported agreement survives when the size term is removed.
The second point concerns the coefficient of determination and the practice of reporting the proportion of variance explained. On data of this shape, that proportion will be very high for the same reason, and reporting it as though it indicated explanatory power is the most misleading single practice in the older studies.
The third point concerns logarithms. Taking logs of both variables converts the multiplicative size term into an additive one, which changes the structure of the problem but does not by itself solve it, since the size term is still present in both variables and still correlated with itself. Some studies have used log transformations and presented them as a control for scale; the transformation reduces the influence of the largest sectors but does not remove the common factor.
The fourth point concerns deviation measures. Rather than correlating, one can compute for each sector the percentage difference between its unit price and its unit value, and then summarise those differences by their mean absolute size, optionally weighting by sector share. This approach measures what a reader actually wants to know, which is how far apart the two vectors are, and it is not vulnerable to the scaling artefact. Its difficulty is normalisation: values and prices are measured in different units, labour time and money, so a conversion factor is required before differences can be computed, and the choice of that factor affects the answer. The usual choice equalises the two vectors in aggregate, which is defensible and not unique.
The fifth point concerns distance measures. Some authors have argued that the appropriate comparison between two vectors is a geometric one, measuring the angle between them, since this is invariant to the scaling of either vector as a whole while still being sensitive to differences in their proportional structure. Measures of this family have been proposed precisely because the correlation statistics were considered flattering, and results reported under them are less impressive than the correlations, which is the expected direction.
The practical rule that follows is short. When a claim about the empirical support for the labour theory of value appears, find out which of these five it rests on. Correlations and coefficients of determination on unnormalised totals carry no information. Deviation measures on per-unit magnitudes carry information whose interpretation depends on the normalisation. Distance measures carry information and are the most conservative. And any of them is uninterpretable without a comparison against alternative candidate bases treated identically.
The critics of the critics
The aggregation critique has itself been criticised, and the exchange is worth reporting because it is where the argument remains live rather than settled.
The first counter-argument is that the critique proves too much. If correlating size-scaled magnitudes is illegitimate, then a great deal of applied economics is illegitimate, since analysis of sectoral data routinely involves quantities scaled by sector size. The response to this is that the objection is not to using such data but to interpreting a correlation between two size-scaled variables as evidence about the relation between their per-unit components, which is a specific inferential error rather than a general prohibition. The counter-argument does not survive, but it does point at something real: the practice being criticised is not unique to this literature.
The second counter-argument is that the random-vector demonstration is less damaging than it appears, because a random vector scaled by size will correlate with price but will not reproduce the structure of the price vector in finer respects. Defenders have argued that labour values outperform random vectors on measures that are less scale-sensitive, and that the demonstration establishes only that the crudest statistic is uninformative. This response is reasonable and it concedes the substance of the critique while contesting its reach.
The third counter-argument is the structural one already described: if near-linear wage-profit curves are observed, the closeness of prices to values follows from the technical structure of production rather than from any statistical artefact, and the aggregation objection does not apply to that finding in the same way. This is the strongest reply available and it moves the dispute onto ground where the critics have engaged less.
The fourth counter-argument is weakest and is worth naming so that readers recognise it. It holds that the critique comes from within a particular interpretive school with its own commitments about temporal valuation, and that its motivation is therefore suspect. Motives are not arguments, the demonstration is reproducible by anyone with the data, and a defender who reaches for this response has run out of better ones.
What the exchange establishes overall is that the first generation of results as originally presented does not stand, that the programme has responded with better measures rather than abandoning the question, and that the number of studies applying the better measures with proper comparative baselines remains small.
Objections to testing the theory at all
A distinct line of criticism holds that the entire enterprise is misconceived, and it comes from inside the Marxist tradition rather than outside it. It deserves a hearing because it is not obviously wrong.
The argument runs as follows. If the labour theory of value is an account of the social form that labour takes under commodity production, rather than a hypothesis about the determinants of relative prices, then computing labour coefficients and correlating them with prices tests a proposition the theory does not make. Worse, the exercise concedes the terms of the opposition by treating the theory as a price-predictive instrument competing with mainstream price theory on mainstream ground, which is precisely the framing that the value-form reading exists to reject. On this view the empirical programme is a category error dressed as rigour.
Two things can be said in response and neither fully disposes of the objection. The first is that the theory as Marx presents it does contain a quantitative claim: socially necessary labour time is a magnitude, prices of production are derived from it by a specified procedure, and the third volume takes the derivation seriously enough to attempt it. A reading on which no quantitative claim is being made has to explain what that derivation is for. The second is that a theory which is protected from all measurement has purchased its security at a price its own adherents may not want to pay, since the tradition’s self-description has generally been scientific rather than hermeneutic.
The stronger version of the objection is narrower and harder to answer. It holds not that measurement is illegitimate but that this particular measurement is testing something too weak to be interesting: that labour requirements track prices tells us little, since labour compensation is a large share of costs in most sectors and any cost-based account would predict something similar. On this version the empirical programme is not a category error but a test with insufficient discriminating power, which is a criticism about design rather than about principle, and it is the one a serious methodologist would press.
How this evidence relates to the other empirical strands
Three other bodies of quantitative work are routinely confused with this one, and separating them prevents most of the misuse of empirical material in argument on this subject.
The first is the measurement of the rate of profit over long periods, which asks whether the profit rate exhibits a downward trend and which turns on how the capital stock is measured, whether financial assets are included, and which sectors are counted. That literature is substantial, its results depend heavily on measurement choices, and it addresses the crisis theory rather than the value theory. It has its own treatment in this series at rate of profit empirical evidence.
The second is the measurement of the rate of exploitation, which computes the ratio of surplus value to variable capital from national accounts, generally using a definitional scheme that makes the calculation tractable. That work tests nothing about the relation between labour and prices; it operationalises a category in order to track it over time.
The third is the wider programme of quantitative Marxian economics, which includes work on the productive and unproductive labour distinction, on the composition of capital, and on the relation between accumulation and employment. The broader survey of that programme belongs to testing Marxist economics, which owns the general question of what in the tradition has been measured.
The rule for keeping these straight is that each addresses a different claim and none substitutes for another. Evidence that the profit rate has fallen says nothing about whether prices track labour values. Evidence that prices track labour values says nothing about the origin of surplus. Citing one as support for another is the commonest form of empirical overreach in this subject, and it happens in advocacy writing on both sides with equal frequency.
The measurement choices that drive the disagreement
In most empirical controversies, the real dispute is about measurement rather than about the phenomenon, and this one is a clear case. Five choices do most of the work.
The choice of labour measure comes first. Using the wage bill as the labour input builds relative wages into the value calculation, which resolves the skilled-labour reduction by assumption and, more damagingly for the exercise, imports price information into the supposedly independent value vector. Using hours worked avoids that but treats an hour of every kind of labour as identical, which the theory itself denies. Neither choice is neutral and results computed with the two are not directly comparable.
The choice of aggregation level comes second, and it is under-discussed. Industries in input-output tables are statistical constructs, and within a broadly defined sector there is enormous heterogeneity. Aggregation tends to average away deviations, so finer disaggregation should be expected to produce weaker agreement, and studies reporting results at coarse levels of aggregation are reporting a partly artefactual smoothness.
The treatment of unproductive sectors comes third. Including finance, trade, and government with everything else, or excluding them, changes both vectors substantially, and the theoretical grounds for the choice are themselves disputed within Marxian economics.
The treatment of fixed capital comes fourth, and it bears most heavily on capital-intensive industries, which are precisely the ones where the theory predicts the largest deviations between values and prices of production. A treatment that handles depreciation crudely will misstate exactly the cases the test should be most sensitive to.
The choice of deviation measure and normalisation comes fifth, and it is where the arguments about magnitude actually live. The same underlying data will yield deviations that look small under one normalisation and substantial under another, and there is no consensus about which normalisation is appropriate.
The practical upshot for a reader is not that the exercise is hopeless. It is that a single reported number means very little, and that the results worth attending to are those that hold across defensible variations in all five choices. Very little of the literature reports that kind of robustness check, which is the clearest indication of how young the field is.
The differential prediction almost nobody tests
There is a sharper test available than the one the literature mostly performs, and its neglect is the clearest sign that the programme has been designed to confirm rather than to discriminate.
The labour theory of value does not predict that prices equal values. It predicts that they diverge, and it predicts the pattern of divergence. Prices of production equalise the profit rate across branches, so a branch with a higher ratio of constant to variable capital than the social average will have a price of production above its value, and a branch with a lower ratio will have a price of production below its value. The magnitude of the deviation should be a function of how far that branch’s capital composition departs from the average.
That is a differential prediction, and it is far more informative than any overall measure of agreement. An overall correlation could arise for many reasons, including the general fact that labour is a large share of cost. The predicted relation between a branch’s capital composition and the sign and size of its price-value deviation could not easily arise by accident, because it is specific to this theory’s mechanism. If deviations were found to be uncorrelated with capital composition, the theory’s account of why prices deviate would be in serious trouble even if the overall agreement were high. If they were found to be correlated in the predicted direction with the predicted sign, that would be evidence of a kind that no arbitrary alternative input measure could produce.
Some studies have examined this relation and reported results consistent with the prediction, and the results are less frequently cited than the headline correlations, which is revealing about the incentives in this literature. The reason it is not the standard test is partly practical: computing the composition of capital by branch requires capital stock data, which is less reliable and less comparable across countries than the flow data in input-output tables. It is also partly rhetorical, since a high correlation is a better advocacy result than a moderately confirmed differential prediction, even though the second is worth much more scientifically.
For anyone assessing this field, the differential prediction is the question to ask about. A study reporting only overall agreement has tested the weaker proposition. A study reporting the relation between capital composition and the direction and size of deviation has tested the theory’s actual mechanism.
Cross-section and time series
Almost all of this literature is cross-sectional: it compares industries with one another at a single moment. The theory also makes a claim about change over time, and testing that claim would be a different and in some ways stronger exercise.
The dynamic proposition is straightforward. If the value of a commodity falls as the labour required to produce it falls, then over a period during which productivity in one branch rises much faster than in another, the relative price of the first branch’s output should fall relative to the second’s, by roughly the amount by which relative labour requirements fell. This is a prediction about changes rather than levels, and it has a property the cross-sectional test lacks: many of the confounding factors that make the level comparison difficult, including the persistent effects of branch-specific market structure, brand position, and regulation, are differenced out when the comparison is between changes over time in the same branch.
The scale-artefact problem also behaves differently in a time series test. Comparing the change in a branch’s unit labour requirement with the change in its relative price does not involve multiplying either quantity by branch size, so the objection that has dominated the cross-sectional dispute does not arise in the same form.
Work of this kind exists and is less developed than the cross-sectional literature, in part because it demands consistent classifications and deflators over long periods, which statistical revisions repeatedly disrupt. It also faces its own difficulty: technical change is often accompanied by quality change, and separating a fall in the price of a computing device from an improvement in what the device does is a measurement problem the national accounts handle by methods that are themselves contested.
The recommendation that follows is one a researcher can act on. The time series test addresses the theory’s dynamic claim, sidesteps the aggregation objection, and has been performed far less. It is the most obvious gap in a literature that has spent thirty years elaborating a single cross-sectional design.
Services, and where the data stops cooperating
The input-output framework was built for an economy of physical commodities flowing between industries, and it handles the sectors that dominate modern output less well. This is not a Marxian problem specifically; it is a general limitation of the data that bears on any cost-based analysis, and it deserves stating plainly because it bounds what any of this evidence can establish.
Measuring output in a service sector requires a unit, and for many services the unit is a convention rather than a physical fact. Output of the financial sector is estimated by methods that have been repeatedly revised and that produce large swings in measured output from changes in interest rate spreads rather than from anything happening in the world. Output of the public sector is conventionally measured by its inputs in many national accounts, which makes any comparison between the labour content of that output and its price close to circular. Output of a legal service, an insurance policy, or a consulting engagement is measured in money and deflated by price indices whose construction is difficult and contested.
The consequences for this literature run in both directions. Studies that include such sectors are correlating labour content with a price measure that in some cases was itself constructed from labour content, which inflates agreement artificially. Studies that exclude them are testing the theory on the shrinking part of the economy where the framework was designed to apply, which is a legitimate restriction but should be stated as one. The distinction between productive and unproductive labour in Marxian accounting overlaps with this problem without being identical to it, and the two are frequently conflated.
A reader assessing any result should therefore establish what share of measured output the study covers and how the difficult sectors were treated. A result covering manufacturing and primary production in a modern economy is a result about a minority of output, which does not make it uninteresting but does make the framing matter.
What the socially necessary clause does to measurement
There is a conceptual difficulty in this exercise that the empirical literature mostly proceeds past, and it is worth stating because it affects how any result should be interpreted.
The theory’s magnitude is socially necessary labour time, not the labour time actually expended. Labour performed under below-average conditions of productivity does not count in full, and labour producing output in excess of what society needs does not count at all. What input-output tables record is labour actually expended in producing output actually produced. The computed vector is therefore a measure of average expended labour by branch, which coincides with socially necessary labour only if the branch is producing an appropriate quantity using representative techniques.
For the first sense of social necessity, the approximation is defensible: an industry average is a reasonable proxy for the prevailing conditions of production in that industry, and the divergence between the two is a within-industry dispersion that aggregate data cannot show anyway. For the second sense, which concerns whether the quantity produced corresponded to social requirements, there is no proxy at all. A branch that overproduced in the period covered has expended labour that the theory does not count, and the table records it.
This means the computed vector is systematically not the theoretical magnitude, and the direction of the error depends on the conjuncture. In a period of general overproduction the computed values overstate socially necessary labour across the board, which mostly cancels in a comparison of relative magnitudes, but in a period where particular branches were badly out of proportion it will not cancel. Studies using benchmark years spanning different phases of the cycle are therefore not strictly comparable.
None of this invalidates the exercise. It does mean that a discrepancy between computed values and prices has two possible sources, a failure of the theory and a failure of the proxy, and that no existing study distinguishes them. That is a genuine limitation and it should temper confidence in results from both directions.
What a decisive replication would require
Because this article’s whole argument is that the literature is thin where it matters, it should specify what would make it thick. A study that did the following would settle a great deal.
It would use recent input-output data for at least three economies with materially different industrial structures, at the finest available level of sectoral disaggregation, and would report results at two or three aggregation levels so the effect of aggregation is visible rather than hidden. It would compute the labour vector in at least two ways, by hours and by an alternative that does not embed wage information, and would report both. It would test at least three candidate value bases besides labour under identical treatment, including at least one randomly generated vector as a null. It would report results under a correlation statistic, a per-unit deviation measure, and a distance measure, so that readers could see how much of the finding is an artefact of the statistic. It would examine the relation between branch capital composition and the sign and magnitude of price-value deviation, which is the theory’s differential prediction. It would state explicitly which sectors were included, how the difficult service sectors were handled, and what share of measured output the analysis covers. And it would be pre-specified, so that the measurement choices were fixed before the results were seen.
None of this is technically demanding. The data is public, the computations are standard linear algebra, and the whole exercise is within reach of a well-supervised doctoral project. That it has not been done is a fact about the sociology of the field rather than about the difficulty of the question, and it is the single most useful thing to know about the state of the evidence.
The honest verdict
The strong version of the empirical claim is not established. That version says the input-output evidence demonstrates that labour time governs relative prices, and it rests on raw cross-sectoral correlations that the aggregation critique shows to be uninformative. No one should cite those correlations as support for the theory, and the fact that they continue to circulate in advocacy writing is a mark against that writing rather than against the theory.
The weak version has some support and remains disputed. That version says labour requirements track relative prices better than chance and better than several alternative input bases once industry size is normalised out. Evidence for it exists, it is not overwhelming, the comparative studies that bear on it are few, and the structural explanations from matrix properties raise a real question about whether the result, if it holds, is specifically confirmatory of anything Marxian.
Confidence should be low in both directions and the reason is not caution for its own sake. It is that the number of studies performing the methodologically correct test, with proper normalisation, with alternative bases treated identically, and with robustness checks across measurement choices, is small enough to count. A field with that few decisive results does not support strong conclusions, and the appropriate description is that the question is open and tractable, which is a more interesting thing to be able to say about a hundred-and-fifty-year-old theory than that it has been proved or refuted.
Which prices are being measured
A technical point that almost never appears in secondary discussion, and that materially affects what the results mean: the price vector in these studies is not the price a consumer pays.
Input-output tables record transactions at producers’ prices, which exclude retail and wholesale margins, transport costs charged separately, and in most treatments the taxes on products. The output of the trade sector appears as its own row, representing the margin, rather than being distributed across the goods it handles. So the comparison being made is between the labour embodied in a unit of a branch’s output and the price that branch receives for it, not between labour content and the shelf price of a finished good.
This is the right choice for testing the theory, because prices of production are producer-level magnitudes and the theory’s mechanism operates through competition between capitals in production rather than through retail markup. But it has three implications a reader should hold on to.
It means the results say nothing about consumer prices, and a claim that these studies show the price of goods in shops reflects labour content is wrong. It means indirect taxes and subsidies, which drive substantial wedges in some categories, are typically outside the comparison, which is analytically correct and removes a source of divergence that a consumer-level test would face. And it means the trade and transport sectors appear as branches with their own labour content and their own prices, which brings them under the productive and unproductive labour question in a way that materially affects the totals.
The broader point is that every one of these decisions is defensible and every one changes the answer somewhat. That is the ordinary condition of applied economics, and the appropriate response is not scepticism about the enterprise but insistence that studies report their choices and test sensitivity to them.
What this literature has actually established
Stripping away the claims made on its behalf, three things have been established by the empirical programme and each is worth having.
The first is that the question is answerable. The labour theory of value generates a measurable implication, the measurement can be performed with public data and standard methods, and the results could have come out otherwise. That closes off the position, held on both sides for most of a century, that the theory is the kind of thing evidence cannot address. Anyone still asserting that it is untestable in principle has an existing literature to explain away.
The second is that price-value deviations in real economies are smaller than the theoretical literature had assumed possible. This is a genuine finding and it stands independently of the aggregation dispute, because it is a claim about magnitudes computed per unit rather than about correlations between totals. The transformation-problem literature had spent decades establishing what could happen; the empirical programme asked what does happen, and the answer was less dramatic than the possibility space suggested. That is a contribution to a debate that had become detached from its object.
The third is methodological and it came from the critics. The demonstration that size-scaled correlations are uninformative is a permanent addition to how this kind of work has to be done, and it disqualifies a substantial part of the earlier literature as evidence. Findings of this sort are how empirical fields improve, and the appropriate response from defenders of the theory is to adopt the standard rather than to dispute it.
What has not been established is the proposition that most people cite this literature for, which is that labour requirements explain relative prices. That would require the comparative and differential tests described above, performed properly and replicated, and the studies doing so are too few to support the claim.
What each side should concede about the evidence
The concessions are asymmetric and specifying them is more useful than another round of the argument.
Advocates should concede that the headline correlations, as originally computed and as still widely cited, do not constitute evidence, because the aggregation objection is correct and is not a matter of interpretation. They should concede that the wage-bill measure of labour input imports price information into the value vector and that studies using it are weaker than they appear. They should concede that coarse aggregation flatters the results. They should concede that the differential prediction about capital composition, which is the theory’s own mechanism, has been tested far less than the overall agreement, and that this is a strange allocation of effort for a programme confident of its case. And they should concede that near-proportionality explicable from the structural properties of input-output matrices is not obviously confirmation of a specifically Marxian proposition.
Critics should concede that the theory generates a testable implication and that the programme is a legitimate scientific exercise rather than an exercise in confirmation. They should concede that demonstrating a test to be uninformative is not demonstrating a hypothesis to be false, and that the arbitrary-input result therefore establishes less than it is usually taken to establish. They should concede that per-unit deviation measures are not vulnerable to the aggregation objection in its simple form, and that results reported under those measures require a different response. They should concede that the wage-profit curve finding, if it holds, is a structural result rather than a statistical artefact and needs to be addressed on its own terms. And they should concede that the measurement difficulties they raise against this literature are shared with orthodox applied work that they do not subject to the same scrutiny.
The residue after both sets of concessions is a well-defined empirical question with an insufficient number of properly designed studies bearing on it, which is a much less exciting description of the situation than either camp offers and is the accurate one.
How this evidence changes the theoretical dispute
The existence of this literature ought to change how the theoretical arguments are conducted, and it mostly has not, which is worth explaining.
The redundancy charge holds that value magnitudes add nothing to the determination of prices and the profit rate. That is a claim about logical necessity within a formal system and empirical results cannot refute it. What empirical results can do is bear on whether the redundancy matters. A magnitude that is logically redundant but empirically tracks the thing it was supposed to explain has a different standing from one that is redundant and also unrelated. Advocates have used the correlations for exactly this purpose, and the aggregation critique undercuts the use; the properly normalised evidence, whatever it turns out to show, is what would determine whether the response works.
The transformation problem is a question about the consistency of a derivation and is not touched by evidence, but the empirical work does affect its significance. If deviations between values and prices of production are small in real economies, then the transformation problem is a real logical difficulty of limited practical consequence, which is a different thing from being a fatal flaw. Both sides have an interest in obscuring this: the critic wants the logical difficulty to be devastating, and the defender wants it to be nonexistent rather than merely unimportant.
The social-form defence is unaffected in one direction and vulnerable in another. Evidence cannot confirm it. But a defender who retreats to it while simultaneously citing the price-value correlations as vindication is holding two incompatible positions, since the correlations are evidence only if the theory makes the quantitative claim the retreat disavows. That inconsistency is common in advocacy writing and it is worth pointing out whenever it appears.
The most useful effect of the empirical programme, in the end, is that it makes it possible to distinguish careful people from careless ones on this subject. Anyone who cites the correlations without mentioning the aggregation objection, and anyone who cites the aggregation objection as though it closed the question, has told you how much of the literature they have read.
Calibrating the numbers against other empirical work
Readers meeting a very high correlation for the first time often find it decisive, and calibration against how such figures behave elsewhere is the fastest corrective.
Correlations of that magnitude are uncommon in individual-level social science data and routine in aggregate economic data. Any two series that both grow with the size of the unit being measured, or both grow with time in a growing economy, will correlate at such levels without any causal relation between them. The classic textbook demonstrations of spurious correlation use exactly this property. The appropriate reaction to a high correlation in aggregate economic data is therefore not surprise but the question of what common factor is driving both series, which is precisely what the aggregation critique asks and answers in this case.
There is a second calibration worth making, and it concerns what a plausible alternative would predict. If one asks what correlation an ordinary cost-based account of prices would generate on this data, the answer is that it would be high too, since labour compensation is a substantial share of costs across most sectors and costs are what prices cover in the long run. The finding is therefore consistent with a proposition much weaker than the labour theory of value, namely that prices reflect costs and labour is a big part of costs. Discriminating between the weak and the strong proposition requires the comparative test, which returns the argument to where it has been throughout.
The third calibration concerns the direction of surprise. Advocates present the results as unexpectedly strong support. Critics present them as unremarkable. Both reactions depend on a prior about what the correlation would be if the theory were false, and neither camp states that prior explicitly. Stating it is a discipline worth adopting: before looking at a result, say what you would expect under the alternative hypothesis. Anyone who does this for the raw correlations will find that they expected a high figure either way, which is another route to the conclusion that the raw figure is uninformative.
What the strongest possible result would and would not prove
Suppose the decisive study described above were carried out and the result came back strongly in favour of labour. Labour requirements track relative prices markedly better than energy, materials, or randomly generated vectors under identical normalisation; the differential prediction about capital composition holds with the right sign; the pattern replicates across three structurally different economies and survives the aggregation levels tested. What would follow?
Considerably less than either camp would claim in the first week. It would establish that the labour input is the best available single-input predictor of relative prices among the candidates tested, which is a substantial and interesting empirical fact. It would not establish that labour is the source of value in the theory’s sense, because a predictive relation is not a causal mechanism and the theory’s mechanism runs through competitive reproduction and the equalisation of the profit rate rather than through any statistical association. It would not touch the redundancy result, which is a claim about logical necessity within a formal system, though it would make redundancy look like a formal curiosity. And it would not bear at all on the origin of surplus or on the social-form claim.
Suppose instead the result came back negatively: labour performs no better than the alternatives after normalisation, and the differential prediction fails. What would follow then?
Again less than expected. It would put the theory’s price claim in serious difficulty and would force defenders onto the social-form reading, which several of them occupy already and regard as the theory’s proper ground. It would not refute the proposition that profit originates in unpaid labour, which does not depend on the price claim and which has been reconstructed independently. It would not vindicate any rival theory of price, since showing that one account fails is not showing that another succeeds.
The exercise of imagining both outcomes is worth performing before looking at any evidence, because it establishes in advance how much any result could matter. The answer here is that the evidence bears on one of the theory’s three jobs and bears on it informatively, which is more than most theoretical disputes in this subject can say and less than the rhetoric surrounding it implies. A reader who holds that proportion in mind will read the studies, the critiques, and the replies without the swings of conviction that this literature produces in people encountering it for the first time.
Three questions to ask anyone citing this evidence
The practical distillation of this article is short enough to remember and it works in both directions.
Ask first whether the compared quantities were per unit or totals. If totals, and no normalisation was applied, the result carries no information about the relation between labour and price, whatever its magnitude and whatever its replication record.
Ask second whether any alternative candidate value basis was tested under identical treatment. Without a baseline, a figure cannot be interpreted, since there is nothing to compare it against. A study that tested energy, materials, and a random vector alongside labour is doing science; a study that tested labour alone is reporting a number.
Ask third whether the study examined the theory’s own differential prediction, which is that price-value deviations should relate systematically to the divergence of a branch’s capital composition from the social average. This is the test that could distinguish the labour theory from any generic cost-based account, and its comparative neglect is the most telling feature of the field.
Three questions, answerable from any paper’s methods section in a few minutes, will place any claim on this subject accurately. That is a better outcome from reading about evidence than a memorised coefficient, and it is the outcome this article is designed to produce.
For students using empirical evidence in an answer
Empirical material earns marks on this topic when it is used to qualify a claim and loses them when it is used to settle one. A candidate who writes that empirical studies have confirmed the labour theory of value has made an assertion a marker can challenge; a candidate who writes that input-output studies report high correlations between labour values and prices, that the correlations are contested because both vectors are scaled by industry size, and that the comparative tests against alternative value bases are the ones that matter, has demonstrated exactly the critical handling that separates upper bands from middle ones.
The distinction that earns marks here is between a theory being untestable and a theory being untested. A great deal of examination writing treats the labour theory of value as a philosophical position immune to evidence, usually in the course of an evaluation paragraph. Saying that the theory has generated a substantial empirical literature, and that the dispute within that literature is statistical rather than philosophical, is a stronger and more accurate move, and it is available to any candidate who remembers one fact about the aggregation problem.
The trap is overclaiming in either direction under time pressure. Naming a specific correlation coefficient from memory is dangerous, since the figures vary by country, year, and method, and a wrong number is worse than no number. The safe formulation names the method, the finding in qualitative terms, and the objection. Answer structures and marks logic for questions on Marx’s value theory belong to the exam and essay guide for Capital Volume One, which owns that guidance; use this article for the evaluation content and that one for the shape of the answer.
For teachers introducing empirical work on this topic
The misconception that dominates a class here is that theories in social science are not the kind of thing that gets tested, which students often absorb from the structure of their courses rather than from anything anyone tells them. Presenting this literature is unusually effective at dislodging it, because the test is comprehensible without technical training: two lists of numbers, one for labour and one for prices, and the question of whether they move together.
The question that surfaces the misconception, and that also teaches the aggregation critique in a single step, is to ask students what they would expect to find if they correlated the total wage bill of each industry with the total sales of each industry. Students see immediately that big industries have big everything, and having seen it, they have understood the objection to the entire first generation of this literature without any statistical vocabulary being introduced.
The extract that resolves the most is a table of industries with output and employment figures for a real economy, which any national statistical agency publishes. Ask a class to eyeball whether the two columns move together, then ask what would happen if the columns were expressed per unit of output instead of as totals. The exercise takes fifteen minutes, teaches spurious correlation, teaches the structure of input-output data, and leaves students able to read the actual dispute rather than being told about it.
For researchers and anyone citing this literature
The verification requirements here are higher than in most of this subject, and three of them are non-negotiable.
Verify the author, the journal, and the year before naming any study. This literature is small, the same handful of researchers appear repeatedly, and a good deal of secondary citation of it is inaccurate about which paper reported which result. Verify whether a cited result is a correlation on aggregate magnitudes or a deviation measure computed per unit, because the two are not comparable and the aggregation critique applies to one and not the other. And verify the measurement scheme, specifically the labour measure, the treatment of unproductive sectors, and the normalisation, because results computed under different schemes are frequently presented in secondary literature as if they were replications of one another.
The claim that is unsafe to repeat is any specific correlation coefficient presented as the finding of the literature. Coefficients vary substantially by country, year, aggregation level, and method, and a figure quoted without its study is meaningless. The safe form is qualitative and ranged: that the reported correlations on aggregate magnitudes are high, that the aggregation critique shows this to be uninformative, and that the normalised comparative results are fewer and weaker.
The reliable material sits in the journal literature rather than in the secondary accounts, and the productive entry point is not a survey but the exchange itself: read a study reporting the correlations, then read the aggregation critique, then read a reply, in that order. Anyone building a file on this should keep each result attached to its dataset, its year, its measurement scheme, and its normalisation, which is the sort of structured note-taking VaultBook’s citation and research tools exist for, and which this topic punishes the absence of more than almost any other in the series.
Has anyone attempted a hostile replication?
Not systematically. The critical work has concentrated on demonstrating that the standard test is uninformative rather than on rerunning the analysis under better specifications to see what survives. That is a real gap, because a critique that stops at showing a method is flawed leaves the substantive question exactly where it was.
The distinction matters for how the whole field should be read. In areas where independent replication by researchers expecting a different result is routine, a body of consistent findings carries substantial weight. In areas where it is absent, consistency across studies may reflect shared methods and shared expectations rather than a robust phenomenon, and the appropriate discount is large. This field is in the second condition. That is not a charge against anyone’s integrity; it is a statement about what a set of results produced under these circumstances can support, and it applies with equal force to the critical side, whose central demonstration has also not been independently reproduced across multiple datasets and specifications by people hoping it would fail.
The remedy is neither complicated nor expensive, which is the frustrating part. A single well-designed adversarial collaboration, in which researchers from both positions agree the specification before seeing the results, would move this question further in one paper than the previous thirty years of parallel argument have managed.
For a researcher looking for a project rather than a summary, the open ground is unusually accessible. The comparative test against alternative value bases, run properly with identical treatment across candidates, on recent tables, at several levels of aggregation, with sensitivity analysis across the five measurement choices, has not been done thoroughly enough for anyone to be confident about the answer. The data is public. The methods are standard. The result would matter to both camps, and a negative finding would be as publishable as a positive one, which is not true of most research questions and is a considerable advantage for anyone choosing a topic. The supervisory requirement is a working knowledge of input-output methods and enough statistical training to specify the normalisation and the sensitivity analysis in advance rather than after the fact, which is where comparable projects in this area have gone wrong.
The vocabulary of these papers
Anyone moving from secondary accounts to the studies themselves meets a specialised vocabulary, and most of it is straightforward once decoded. Knowing these eight terms makes the primary literature readable.
The technical coefficients matrix records how much of each industry’s output is required as input per unit of each other industry’s output. It is the raw material of the whole exercise and is constructed directly from the input-output table.
The Leontief inverse is what you get by inverting one minus that matrix. It converts a vector of final demands into the total gross output required to satisfy them, counting every round of indirect requirement. Multiplying it by the vector of direct labour inputs gives the total labour requirement per unit of final output.
The vertically integrated labour coefficient is that resulting quantity for a single industry: all the labour, in that industry and in every industry supplying it and every industry supplying those, required to produce one unit. This is the empirical stand-in for labour value, and the phrase vertically integrated signals that the whole chain has been collapsed into one number.
The monetary expression of labour time is the conversion factor between labour hours and money, needed because values are measured in time and prices in currency. Different definitions of it exist, and the choice is not innocent: some definitions make certain aggregate identities hold by construction, which is why an author’s definition should be checked before their identities are taken as findings.
The numeraire is the good or bundle in terms of which everything else is priced. Results in this literature can be sensitive to the choice, and a study that does not state its numeraire has omitted something material.
The standard commodity is a construct from the classical revival, a composite good whose composition makes its price invariant to changes in distribution. It appears in this literature as a benchmark for normalisation and as a device for isolating the effects of distributional change from the effects of technique.
The organic composition of capital is the ratio of constant to variable capital, adjusted so that it reflects the technical relation rather than price movements. It is the variable that the theory says should predict the direction and magnitude of price-value deviation, and measuring it requires capital stock data.
Mean absolute deviation and its weighted variant are the per-unit distance statistics, computed by taking the percentage gap between the price and the value of each unit of output and averaging across industries, weighting by industry share in the weighted version. These are the measures that are not vulnerable to the aggregation objection in its simple form.
The planning connection
One motivation behind part of this literature is rarely mentioned in secondary accounts and explains some of its features.
Several of the researchers who developed the price-value correlation work were also interested in whether labour time could serve as an accounting unit for a planned economy, on the argument that if computing labour coefficients for a whole economy is feasible with modern data and computing power, then the informational objection to planning based on labour time is weaker than it was when the calculation debate began. On this view the empirical work is not only a test of a descriptive theory but a feasibility study for an alternative accounting system.
That motivation is not a criticism, but it is context a reader should have. It explains why the emphasis fell on demonstrating that labour coefficients can be computed for real economies at scale, which is a different objective from discriminating between competing explanations of relative prices, and it explains why the comparative and differential tests received less attention than the demonstration of feasibility. Research programmes are shaped by what their practitioners want to establish, and the shape of this one has consequences for what it has and has not tested.
The connection also matters for how the results should be used in argument. A finding that labour coefficients can be computed accurately and that they track prices reasonably well bears on the feasibility of labour-time accounting somewhat independently of whether the labour theory of value is the correct account of price formation under capitalism. Those are separate propositions and the literature sometimes runs them together. A reader interested in the planning question and a reader interested in the value theory are, strictly, interested in different implications of the same computations.
Why confidence should stay low
It is worth being explicit about why the verdict here is a split rather than a lean, because readers are entitled to suspect that a split verdict is a way of avoiding a conclusion.
The reason is not that the evidence is balanced. It is that the evidence bearing on the question, properly measured, is scarce. A large number of studies exist; a small number of them perform the test that would be informative. Counting papers is not counting evidence, and a literature of forty studies of which four are decisive is a literature of four studies with a lot of context.
The second reason is that almost all of the work comes from researchers who expect one answer. This is not an accusation of bad faith and the individual studies are not thereby wrong. It is a structural observation about how much weight a body of results can bear when independent replication by people who expect a different outcome is largely absent. In fields where this condition has been studied, results have generally not held up as well as their original authors expected, and there is no reason to think this field is exempt.
The third reason is the number of researcher choices between the raw tables and the reported statistic. Five decisions, each defensible in more than one direction, means a large space of possible analyses, and a literature that reports the analysis performed rather than the space explored is not in a position to claim robustness. The remedy is pre-specification and sensitivity reporting, which almost no study in this area has practised.
A reader who wants certainty from this literature will not get it. A reader who wants to know what would produce certainty can have that, which is the more valuable possession and is the reason for setting out the design requirements above in as much detail as this article has.
The state of the evidence
The most useful thing this literature has produced is not a correlation but a demonstration that the question is answerable. For most of its history the labour theory of value was argued about as though evidence were unavailable in principle, and both camps found the arrangement comfortable. The input-output programme ended that, and the aggregation critique then showed that the first generation of results did not mean what was claimed. Both of those are advances.
What remains is a well-defined question with a clear method and an insufficient number of studies. Labour requirements may track relative prices better than alternatives once size effects are removed, or they may not, and the work that would settle it is neither conceptually difficult nor expensive. The reason it has not been done at scale is that the small number of researchers with the technical capacity and the interest are mostly committed to one side, and confirmation from a committed source is worth less than a hostile replication, which is a general problem in empirical social science and an acute one here.
Anyone who wants a verdict rather than a research agenda should take the split one: the strong claim is unestablished, the weak claim is disputed with some support, and the confident statements circulating on both sides of this argument are not supported by the literature they cite.
That verdict will disappoint readers who arrived wanting to know whether Marx was right. What it offers instead is more durable. A theory that has generated a measurable implication, attracted a research programme, produced results, had those results substantially undercut by a methodological objection, and responded by improving its measures is a theory being treated as a scientific proposition rather than as an article of faith or an object of ridicule. That is an unusual condition for anything in this subject, and it is worth more than a verdict that the evidence cannot yet support.
Frequently Asked Questions
Q: Is there empirical evidence for the labour theory of value?
There is a real empirical literature and its results are contested. Studies using national input-output tables compute the labour required directly and indirectly to produce each industry’s output and compare it with observed prices, and they report high correlations across many countries. The objection is that both quantities are totals scaled by industry size, so they will correlate strongly whatever the relation between the underlying per-unit magnitudes, and randomly generated vectors treated the same way produce comparable results. Evidence that survives normalisation exists but is thinner. The honest summary is that the strong claim is unestablished and a weaker claim has some disputed support.
Q: How do you measure labour values in practice?
From published input-output tables. These record what each industry buys from every other industry and what it pays in wages over a period. Inverting the matrix of technical coefficients and multiplying by direct labour inputs gives the total labour, direct plus indirect, required per unit of each industry’s output. Several choices arise along the way: whether to measure labour as hours, persons, or wage bill; how to treat fixed capital; whether to include sectors Marxian accounting classifies as unproductive; how to handle imports; and how finely to disaggregate industries. Each choice moves the results, which is why single reported figures should be treated cautiously.
Q: What do the price value correlation studies show?
Consistently high correlations between the labour content of industry output and the money value of that output, reported across the United States, the United Kingdom, Greece, Mexico, Sweden, Yugoslavia, Japan, Italy, and others. Authors describe the agreement as strong and considerably closer than the theoretical transformation literature had led people to expect. The finding held up across replications, which advocates present as robustness. The difficulty is that the standard version of the test compares aggregate magnitudes rather than per-unit ones, which makes the correlation partly a measurement of how industry sizes are distributed rather than a finding about the relation between labour and price.
Q: Why do critics say those correlations are spurious?
Because both variables are scaled by the same third variable. The labour value of an industry’s output equals labour per unit times units produced; the price of that output equals price per unit times units produced. Industry sizes in a national economy vary by orders of magnitude, and two quantities both multiplied by a widely varying common factor will correlate strongly regardless of any relation between the underlying per-unit figures. The demonstration that settles the point is empirical: randomly generated vectors and arbitrary non-labour inputs, scaled the same way, produce correlations of the same order. That shows the raw test is uninformative, though it does not show the underlying hypothesis is false.
Q: Which countries have been tested for price value correlation?
The programme has been run on the United States and the United Kingdom most extensively, and replications or extensions exist for Greece, Mexico, Sweden, Italy, Japan, the former Yugoslavia, and several others, along with multi-country comparisons attempting to establish whether the pattern holds across different industrial structures and levels of development. The consistency across such varied economies is presented by advocates as evidence against an artefactual explanation. The counter-argument is that if the result is driven by the size distribution of industries, and that distribution is broadly similar across market economies, then cross-country replication is exactly what an artefact would produce.
Q: Can the labour theory of value be tested at all?
Yes, and the claim that it cannot is refuted by the existence of the literature. What is testable is the proposition that labour requirements govern relative prices of reproducible commodities, which yields a measurable prediction: labour content should track prices more closely than alternative input measures do. What is not testable by this method is the claim about the origin of surplus, or the claim about the social form that labour allocation takes, both of which are propositions of a different kind. Anyone asserting that the theory is unfalsifiable in principle should be asked to explain what the input-output studies have been doing for a generation.
Q: What data do you need to test Marxist value theory?
A national input-output table with sufficient sectoral detail, employment or hours data by sector, and output values by sector, all for the same period and on consistent definitions. Statistical agencies in most industrialised countries publish all three. Fixed capital data is needed for anything beyond the simplest treatment. The technical requirement is matrix inversion, which is trivial computationally. The genuine difficulties are definitional rather than computational: deciding what counts as productive labour, how to treat imports embodying foreign production conditions, and at what level of aggregation to work. Those decisions do more to determine results than any statistical technique applied afterwards.
Q: Do the studies control for industry size?
The first generation largely did not, which is the whole basis of the aggregation critique. Later work has used per-unit measures, deflated magnitudes, and alternative normalisations that remove or reduce the size effect, and results computed this way are weaker than the raw correlations, as expected. The most informative approach compares labour against alternative candidate value bases under identical scaling, so that whatever artefact affects one affects all. Studies doing this properly are a small minority of the literature. When assessing any claim about empirical support, the first question to ask is whether size was normalised out, because unnormalised results carry no information.
Q: Have the correlation results been replicated?
Repeatedly, across many national datasets and by several research groups, and the replications report results of similar magnitude. Whether replication constitutes support depends on what is driving the result. If the correlation reflects a real relation between labour and price, replication across diverse economies is strong evidence. If it reflects the size distribution of industries, replication is equally expected, since that distribution does not differ dramatically between market economies. The replications therefore do not discriminate between the two explanations, which is why the comparative tests against alternative value bases, rather than further replications, are the studies that would advance the question.
Q: What would a genuine empirical refutation look like?
A properly normalised comparison, on good data at several levels of aggregation, showing that labour requirements track relative prices no better than arbitrary alternative input measures, and that the result holds across defensible variations in the labour measure, the treatment of unproductive sectors, the handling of fixed capital, and the choice of normalisation. That would leave the theory’s price claim without empirical content and force its defenders entirely onto the social-form reading. Note what would not constitute refutation: demonstrating that the raw correlations are uninformative, which has been done, shows only that a particular test tells us nothing, not that the hypothesis it tested is false.
Q: Do these studies settle the transformation problem?
No, and the confusion between the two questions is one of the most common errors in secondary writing on this subject. The transformation problem concerns whether Marx’s procedure for converting values into prices of production is internally consistent, and whether both aggregate identities can be preserved simultaneously under a given valuation convention. That is a question about the logic of a derivation, and no empirical correlation bears on it. Findings of close agreement between values and prices do not resolve it; findings of large divergence do not create it. The two disputes have run in parallel for decades with very little genuine contact between them.
Q: Why do high price value correlations not prove Marx right against mainstream economics?
Because the finding is compatible with both frameworks. If long-run competitive prices approximate marginal costs, and labour compensation is a large share of value added throughout an economy, then a close relation between labour requirements and prices is roughly what a mainstream account would also predict. Establishing that one framework fits the data does not establish that a rival does not, and the discriminating test, one whose outcomes the two frameworks predict differently, has not been constructed. Presenting the correlations as a victory over marginalism therefore overstates what they show, even before the aggregation critique is considered.
Q: What is the strongest evidence available for the labour theory of value?
Not the headline correlations, which do not survive the aggregation objection. The strongest available evidence is the comparative work testing labour against alternative candidate value bases under identical treatment, since that design controls for whatever artefact affects all candidates equally. Alongside it sits the theoretical work explaining why near-proportionality between prices and labour values should follow from the structure of real input-output matrices, particularly the dominance of their first eigenvalue. That line is the most interesting in the area and it is double-edged: if near-proportionality follows from a general structural property, it may not be confirmation of anything specifically Marxian.
Q: How large are the deviations between prices and labour values in real economies?
Reported deviations depend heavily on the measure and the normalisation, which is itself the substance of a methodological dispute, so no single figure should be quoted as the answer. Studies using mean absolute or mean percentage deviation on per-unit magnitudes generally report deviations that their authors describe as modest and that critics describe as substantial, with the disagreement turning on what baseline counts as modest. What can be said without qualification is that the deviations found in real economies are smaller than the theoretical literature on the transformation problem had suggested was possible, which was the original point the empirical programme was making.
Q: Is the input-output method biased toward finding a correlation?
There are reasons to think it is, and honest defenders acknowledge them. Using the wage bill as the labour measure imports price information directly into the value vector, which builds in agreement. Coarse aggregation averages away deviations, so results at broad sectoral levels look smoother than the underlying reality. Treating depreciation crudely misstates precisely the capital-intensive industries where the theory itself predicts the largest divergence. None of these makes the exercise worthless; all of them mean that a study should report sensitivity to these choices, and that a study which does not report such sensitivity should be read as preliminary rather than as a result.
Q: What should a journalist say in one sentence about evidence for Marx’s value theory?
That labour content and prices track each other closely across industries in the published data, but that the standard measurement of this relation is contested because both quantities scale with industry size, and the tests that control for this are fewer and less conclusive. That sentence is defensible and will survive contact with a specialist from either camp. The two formulations to avoid are that the theory has been empirically confirmed, which overstates what survives the aggregation critique, and that it is untestable in principle, which is contradicted by a generation of published work.
Q: Where should a postgraduate looking for a project in this area start?
With the comparative test done properly, because it has not been. Take recent input-output tables for several economies, compute labour requirements alongside two or three alternative candidate bases, apply identical normalisation to all of them, and run the comparison at several levels of aggregation with sensitivity analysis across the labour measure, the productive and unproductive boundary, and the treatment of fixed capital. The data is public, the methods are standard, and the result would be informative whichever way it came out. A hostile replication would be more valuable still, since almost all existing work in this area comes from researchers committed to one answer.
Q: Why is this literature so hard to find outside academic journals?
Partly because it is technical enough to resist compression, and partly because it does not deliver what either camp wants. The correlations look like a vindication until the aggregation critique is applied, so the results are unusable in advocacy. The critique looks like a refutation until one notices that it shows a test is uninformative rather than that a hypothesis is false, so it is unusable in the other direction. Material that supports no confident conclusion travels poorly, and the result is that the most genuinely empirical work on this subject is nearly invisible to anyone not reading the specialist journals directly.