By Giorgia Lazzarotto
1. The analytical problem
The relationship between privacy, artificial intelligence, and decision-making is usually debated as a matter of law or ethics. This article approaches it as a measurement problem, using the apparatus of computational social science (CSS). CSS is the study of social behaviour through large-scale digital data, machine learning, simulation, and network analysis. Its founding programme (Lazer and colleagues, 2009, updated 2020) argued that digital traces let us observe social life at a granularity that was previously impossible. That same granularity is exactly what makes contemporary decision-making systems possible, and it is where privacy is won or lost.
Three shifts make the young generations a distinct case rather than a smaller version of the adult one. First, people born after roughly 2000 generate a near-continuous behavioural record from early childhood, so the training data that will later govern them already exists before they can meaningfully consent to it. Second, many decisions that used to be discretionary and human – which welfare service to recommend, which applicant to fast-track, which student to flag as at risk – are increasingly delegated to statistical models. Third, the young are the population on which the state most actively intervenes for developmental and protective reasons, which means their data is not only abundant but institutionally acted upon.
Finland is an unusually clean site to study this. It combines a comprehensive, register-based welfare state (so the data exists and is linkable), high public trust and digital adoption, and an unusually strong tradition of legal scrutiny over administrative automation. The result is a natural experiment: a society that built ambitious AI-in-government infrastructure and then, in full public view, tested it against its own constitutional and privacy commitments. The Finnish record from roughly 2018 to 2024 gives us concrete cases to analyse rather than hypotheticals.
2. A computational-social-science apparatus
To analyse privacy, AI, and youth decision-making rigorously, it helps to be explicit about the methods involved, because each method carries its own privacy signature.
Digital trace and behavioural data. The raw material of CSS is the exhaust of ordinary life: administrative registers, service-use logs, sensor and app data. Matthew Salganik’s framing (Bit by Bit, 2018) is useful here – trace data is big, always-on, and non-reactive, but also incomplete, non-representative, and drifting. For young people the trace is denser and longer than for any prior cohort, which improves prediction and simultaneously raises the stakes of re-identification.
Supervised prediction and profiling. Most “AI” in public services is supervised machine learning: a model learns the statistical relationship between features and an outcome, then scores new individuals. Kleinberg and colleagues (prediction policy problems, 2015) showed why governments find this attractive – many policy tasks are, at their core, prediction tasks. The privacy issue is that accurate prediction rewards data maximisation, which is in direct tension with data minimisation, the core principle of European data-protection law.
Automated decision-making (ADM). When a prediction is wired directly to an administrative outcome without a human deciding, prediction becomes decision. This is where CSS meets due process, and where Finland’s cases are sharpest.
Microsimulation and agent-based modelling. Not all computational methods profile individuals. Microsimulation models such as Finland’s SISU simulate how a change in tax and benefit law would affect a representative population, using register microdata to estimate distributional effects before a policy is enacted (Statistics Finland, SISU). This is a decision-support use of computation that is designed to be privacy-preserving: it reasons over a synthetic or protected population, not over an identified citizen. The contrast between SISU-style simulation and AuroraAI-style individual profiling is one of the central distinctions this article draws.
Causal inference versus prediction. A recurring error in applied CSS is to treat a predictive correlation as if it were an actionable cause. A model that predicts which young person will drop out is not a model of why they drop out, and intervening on the predicted individual can be both ineffective and stigmatising. This distinction matters for youth policy, where the temptation to act early on a risk score is strongest.
Algorithm auditing and accountability. Finally, CSS includes the study of algorithms as objects of governance: auditing models for bias, demanding transparency, and testing them against legal norms. In the Finnish case the auditor is not a research team but a constitutional institution, the Parliamentary Ombudsman, which gives us a rare public audit of state ADM.
Each of these methods implies a different privacy model: differential privacy and pseudonymisation for protected computation, contextual integrity (Helen Nissenbaum) for judging when a data flow is appropriate, and networked privacy (danah boyd) for understanding why individual consent is a weak instrument in connected populations. I return to these below.
3. The datafication of young people’s decisions
Before turning to Finland, it is worth stating precisely what is being decided and by whom. Young people’s decisions are datafied at three levels.
At the individual level, recommendation and scoring systems shape the choice set a young person sees: which course, which benefit, which service, which opportunity is surfaced or hidden. Sonia Livingstone’s work on children’s data rights makes the point that this is not only a privacy question but an agency question, because the architecture of choice is set upstream of the choice itself.
At the institutional level, agencies use predictive models to allocate scarce attention: whom to contact, whom to prioritise, whom to flag. Virginia Eubanks (Automating Inequality, 2018) documented how such systems, even when well intended, tend to concentrate scrutiny on the already-visible poor, because they are the population whose data the state holds most completely. Young people in contact with welfare, health, or child-protection systems are precisely such a high-visibility population.
At the societal level, simulation and forecasting inform the design of the rules themselves – the tax schedules, benefit tapers, and eligibility thresholds that structure a generation’s economic decisions. This is the SISU register.
The privacy stakes differ at each level. Individual scoring risks profiling and self-fulfilling classification. Institutional allocation risks feedback loops and unequal exposure. Societal simulation, done well, is comparatively benign, because it does not require acting on an identified person. A serious analysis therefore cannot say “AI and privacy” in the singular; it has to ask which computational method is doing which work.
4. Finland as a natural laboratory
Finland is a maximally datafied welfare state. Its administration runs on linkable personal registers, and automated processing is not a future prospect but an established reality: the Deputy Parliamentary Ombudsman found in 2019 that more than 80 percent of the Tax Administration’s decisions were fully automated (AlgorithmWatch, Automating Society Finland, 2020). At the same time, Finland pursued an explicit strategy to make AI a civic competence rather than an elite one, through the free Elements of AI course launched by the University of Helsinki and Reaktor in 2018, whose stated ambition was to teach the basics of AI to one percent of the population and which was later offered across the EU.
What makes Finland analytically valuable is that this ambition met resistance from within the state’s own legal order. The same country that automated most tax decisions also produced the clearest public findings that such automation lacked adequate legal basis. This internal tension, rather than any single project, is the real object of study.
5. AuroraAI: prediction, profiling, and the youth digital twin
The clearest expression of the individual-scoring paradigm was AuroraAI, Finland’s national AI programme for human-centric public services, launched as a national programme in 2020 and coordinated by the Ministry of Finance (Finnish Government, 2020). Its animating idea was the “life event”: at moments such as leaving school, having a child, or recovering from injury, AuroraAI would connect a person to relevant services across government, private, and community providers, using AI to route the individual to the right help.
Read through the CSS apparatus, AuroraAI was an individual-level recommendation and profiling system. To route a person well, it had to infer that person’s situation and likely needs from data, which is supervised prediction applied to life circumstances. The critiques that accumulated map precisely onto the method’s known failure modes (The Mandarin, 2023). First, autonomy: a system that infers a person’s goals and nudges them toward AI-determined service pathways substitutes a modelled preference for a stated one. Second, bias and drift: inferences trained on historical data risk reproducing the social patterns encoded in that data. Third, and decisively in practice, data aggregation: assembling a coherent picture of a person from fragmented municipal, national, private, and community sources ran into exactly the legal barriers that European data-protection and Finnish administrative law erect against unconstrained linkage.
The youth dimension is not incidental. In 2022 AuroraAI was piloted in Zekki, a self-assessment and service platform for young people, and the reported result was telling: the recommendations were “useful but generic” rather than genuinely personalised. From a CSS standpoint this is the signature of the cold-start and sparsity problem – for a young person the informative, longitudinal trace either does not yet exist or is legally walled off, so the model falls back on population averages. The privacy-protective legal environment and the technical limits therefore pointed the same way. AuroraAI was quietly wound down, and maintenance of the network’s core components ended on 31 December 2023 (Digital and Population Data Services Agency, 2023).
The instructive reading is not that the project failed technically. It is that an individual-profiling approach to youth decision-making collided with two independent constraints at once: a data-protection order that resists the linkage such profiling needs, and the developmental reality that the young are precisely the group for whom accurate individual prediction is hardest and most ethically fraught. When the profile is thin, the system either underperforms (generic advice) or overreaches (inferring what it cannot know). Both outcomes are visible in the Finnish record.
6. Automated decision-making and the due-process audit
If AuroraAI shows the limits of recommendation, the automated-decision-making cases show what happens when prediction is fused to a binding outcome. Here Finland offers something rare: a public, quasi-judicial audit of state algorithms conducted by a constitutional institution.
In November 2019 the Deputy Parliamentary Ombudsman found the Tax Administration’s automated decision-making unlawful, on three grounds that read like a due-process checklist for ADM (AlgorithmWatch, 2020). The first was legal basis: automation was resting on general legislation that never explicitly authorised deciding by algorithm, and the absence of a defined, inspectable procedure defeated public scrutiny. The second was accountability: responsibility had been assigned to abstract “process owners” rather than to identifiable officials, which conflicted with the Finnish constitutional and criminal-law requirement that a named person answer for an administrative act. The third was good governance: citizens were not told when a decision was automated and could not obtain the basis of the decision. In parallel, the Chancellor of Justice opened an investigation into the Social Insurance Institution, Kela, which settles on the order of 15.5 billion euro of benefits annually, over the opacity of which of its decisions were automated. The Immigration Service, Migri, met objections from Parliament’s Constitutional Law Committee when it sought to automate permit processing.
The CSS lesson is that these are not primarily complaints about model accuracy. They are complaints about the collapse of the distinction between prediction and decision without a corresponding legal and procedural apparatus. An accurate model attached to an unaccountable process is still unlawful. Finland’s response was legislative rather than merely critical: general legislation on automated decision-making in public administration, together with amendments to the Administrative Procedure Act, entered into force in 2023, providing an explicit legal basis for ADM alongside safeguards on documentation, transparency, and accountability (Finnish Government, 2023).
For young people the significance is concrete. The decisions most likely to be automated at scale – student benefits, housing support, unemployment and activation measures – fall heavily on people at the start of adult life, who are least equipped to detect and contest an erroneous automated outcome. A right to know that a decision was automated, and to obtain its basis, is in practice a right that protects the young disproportionately, because they are over-represented in exactly the high-volume, low-discretion decision streams that automation targets first.
7. The human-centric counter-model: MyData and Findata
Finland did not only produce a cautionary tale. It also produced two of the more coherent institutional answers to the privacy-versus-utility dilemma, and both are analytically important because they change where control and computation sit.
MyData is a model of human-centric personal data management that originated in Finland, articulated in a 2015 white paper by Poikola, Kuikkaniemi, and Honko and subsequently carried by MyData Global, a Helsinki-based organisation, through its declaration of principles (University of Helsinki research portal; MyData Declaration). Its core claim is that individuals should hold actionable control over data about them, including rights of access, portability, and the ability to authorise flows to third parties. In CSS terms, MyData tries to relocate the point of consent from a one-off checkbox to an ongoing, revocable, machine-readable permission, which is a direct response to the well-documented failure of notice-and-consent at scale.
Findata, the Finnish Social and Health Data Permit Authority established in 2020 under the Act on the Secondary Use of Health and Social Data (in force from 2019), addresses the other half of the problem: how to allow computational research and development on sensitive data without exposing identified individuals (Findata). The mechanism is a permit-based system in which linked health and social data are pseudonymised and analysed inside a secure processing environment, so that researchers obtain statistical results rather than raw identities. This is, in effect, the institutionalisation of privacy-preserving computation: it lets society reason over its young population’s health and social outcomes at the aggregate, model-building level while structurally denying analysts the individual re-identification that individual profiling requires.
Set side by side, MyData, Findata, and SISU form a coherent alternative to the AuroraAI paradigm. Where AuroraAI sought to compute over the identified individual to steer that individual’s decisions, this second family computes either with the individual’s revocable authorisation (MyData) or over a protected, de-identified population (Findata, SISU) to inform collective decisions. The privacy difference is not one of degree but of architecture. It is the difference between contextual integrity preserved and contextual integrity breached, in Nissenbaum’s sense: data collected for care or administration, flowing on to statistical use under controlled and legible norms, versus data collected in one context being repurposed to profile and nudge in another.
8. Youth-specific dynamics: consent, the privacy calculus, and networked autonomy
Why treat the young as a distinct case rather than as ordinary data subjects? Three findings from the behavioural and computational literature justify it.
First, the so-called privacy paradox is strongest among young users: stated privacy concern predicts disclosure behaviour weakly, because disclosure is driven by immediate social reward under conditions of bounded rationality (the privacy calculus). A model trained on youth behaviour therefore learns from choices made under systematically distorted incentives, and treating those choices as revealed, autonomous preferences launders a structural condition into an individual signal.
Second, privacy for the young is networked, not individual (danah boyd). A teenager’s data is co-produced by peers, parents, schools, and platforms, so individual consent cannot in principle protect it; the relevant unit is the network, not the person. This is exactly the property that makes profiling feasible – the missing data about one young person can be inferred from their connected others – and exactly the property that individual-rights instruments fail to capture. Finland’s welfare registers are a formal version of this networked substrate, which is why the legal question kept returning to linkage and legal basis rather than to any single consent.
Third, developmental autonomy is a moving target. The young are, by design, in the business of becoming different from their past selves. A system that predicts a young person from their historical trace and then acts on that prediction risks freezing an identity that the person is entitled to outgrow, a concern that maps onto the European right to erasure but goes beyond it. The AuroraAI intuition of inferring a person’s goals is most problematic precisely for those whose goals are not yet settled.
These three findings converge on a single methodological warning. For adults, one can at least argue that behavioural data reflects settled preferences under normal incentives. For the young, the data reflects preferences that are unsettled, socially co-produced, and formed under reward structures optimised by the very platforms doing the measuring. The inferential ground is therefore weakest exactly where the interventionist temptation is strongest.
9. What computational social science can and cannot see
An honest analysis must turn the lens on itself, because CSS is not a neutral observer of this system; it is one of its instruments.
CSS is strong at description and prediction. It can show, from register and trace data, how young people’s outcomes cluster, how service use flows through a population, and how a benefit reform would redistribute income (the SISU use case). These are legitimate and valuable, and notably they are the uses that can be done under Findata-style protection without individual profiling.
CSS is weak, and dangerous, when predictive accuracy is mistaken for causal understanding or for a mandate to act on individuals. A model that predicts youth disengagement with high accuracy still does not explain it, and deploying it to target individuals converts a descriptive instrument into a decision that demands the due-process apparatus the Finnish Ombudsman insisted on. The methodological distinction between prediction and causation is not academic hygiene here; it is the line between analysis and unaccountable governance.
There is also a reflexive privacy cost specific to measurement. To study privacy behaviour computationally, one must collect behavioural data, so the research itself sits inside the phenomenon it studies. Finland’s answer – permit-gated, pseudonymised analysis in secure environments – is as much a model for responsible CSS as it is for responsible administration. The same architecture that protects the citizen from the state’s profiling protects the research subject from the researcher’s.
Finally, CSS should be candid about representativeness. Register data captures the population in contact with institutions, which over-represents the visible and the governed. Building youth models on that base risks Eubanks’s feedback loop: the young people the system can see most clearly become the young people it scrutinises most heavily, and the resulting data confirms the original focus. No amount of model sophistication corrects a sampling frame that is itself a product of unequal institutional attention.
10. Conclusion
Analysed through the methods of computational social science, the privacy-AI-decision-making problem for young generations resolves into a single distinction with large consequences: computing over the identified individual to steer that individual, versus computing with revocable authorisation or over a protected population to inform collective choices. Finland ran both experiments in public. The first, AuroraAI, met the limits of profiling a population whose data is thin, networked, and legally guarded, and was wound down. The second – the Ombudsman’s insistence on legal basis and accountability, the 2023 automated-decision-making legislation, and the MyData and Findata architectures – shows what it looks like to keep the analytic power of computation while denying it the individual re-identification and the unaccountable decisioning that make it hazardous.
The lesson is not that the young should be excluded from datafied services, which is neither possible nor desirable, but that the method must be matched to the stakes. Simulation and de-identified analysis can and should inform the rules that shape a generation’s choices. Individual profiling and automated decision-making, where they touch the young, must carry the full weight of transparency, contestability, and named accountability, because the young are the population least able to detect an error and most entitled to become someone their past data did not predict. Computational social science supplies both the tools that create this tension and the discipline needed to hold it. Finland’s decade of trying is the closest thing we have to a controlled test of whether a society can have the first without surrendering to the second.
Sources
- Finnish Government, The AuroraAI national artificial intelligence programme begins (2020)
- Digital and Population Data Services Agency, Maintenance of the AuroraAI network will end on 31 December 2023
- The Mandarin, What Australia can learn from Finland’s AI disaster (2023)
- AlgorithmWatch, Automating Society Report 2020 – Finland
- Finnish Government, Digital and Population Data Services Agency in 2023: automated decision-making and new digital services
- Findata, Finnish Social and Health Data Permit Authority and legislation on secondary use
- MyData, About the MyData Declaration; University of Helsinki, MyData: an Introduction to Human-Centric Use of Personal Data
- Statistics Finland, SISU microsimulation model
Note: the analytical framings referenced in the text draw on established scholarship in computational social science and data ethics – among them Lazer and colleagues on computational social science, Salganik’s Bit by Bit, Kleinberg and colleagues on prediction policy problems, Nissenbaum’s contextual integrity, danah boyd on networked privacy, Sonia Livingstone on children’s data rights, and Virginia Eubanks’s Automating Inequality. These are cited at the level of concept and author; readers should consult the primary works for exact formulations.
This is an analytical essay for discussion; it is not legal advice.
Leggi anche
- Datafizierte Jugend, automatisierte Entscheidung: Privatheit, KI und die Entscheidungsfindung junger Generationen im Licht der Computational Social Science – der finnische Fall 22 Lug 2026
- Gioventù datificata, scelta automatizzata: privacy, IA e decisioni delle nuove generazioni attraverso le scienze sociali computazionali – il caso finlandese 22 Lug 2026
- Consulenza compliance a distanza: cosa funziona da remoto e cosa no 22 Lug 2026
