Abstract
The GDPR impacts the design of information systems which process personal data, because it makes mandatory the adoption of the privacy-by-design and privacy-by-default principles. This compliance must be verified throughout the design cycle, so that it must be considered as early as possible in the cycle, when alternatives are not yet detailed in the overall design and just general directions of the projects may be available. A comparison between alternatives should be performed, which can only have a qualitative nature, but which involves numerous factors, so a panel of experts is needed to obtain a reliable result. In this paper, we propose a analytic hierarchy process-based evaluation approach to examine privacy-related features of alternative information system architectures in the early phases of the design cycle.
Keywords
Introduction
Privacy-by-design and privacy-by-default are mandatory requirements imposed by the GDPR to all systems dealing with personal data operating or produced, or sold in the EU. They are non-functional requirements which generate a number of functional requirements or modify functional requirements to ensure compliance and reduce risk. 1 While the adoption of these design principles in design processes of software systems may be managed by approaches which allow, at some cost, fixes and countermeasures in consequences of alarming results of intermediate or continuous assessments, when designing (or redesigning, or maintaining) complex hardware/software (HW/SW) systems the decisions about the architecture of the system are not reversible and heavily affect all the other aspects of the design and costs, and privacy-by-design and privacy-by-default must be applied since the preliminary decisions about the future system or the evolution of an existing system. This is especially impacting large and complex distributed systems. It is also essential to take into account other GDPR provisions when making architectural decisions, including Article 25 (processing principles), Article 6 (lawfulness), Article 17 (right to erasure), Article 20 (data portability), and Articles 32–34 (security of processing, breach notification, and communication).
An example is provided by the smart city domain. In this field, off-the-shelf components are largely offered on the market and the benefits, features and cost model provided by the Cloud/Edge paradigm can be exploited with no big exposure to financial or technological risk, as there is also a consolidated practice for data collection procedures which have a modest impact on running costs, suitable for unattended installation over large areas and intelligent preprocessing of data streams and complex information. This does not hold, unfortunately, for privacy risk (and security risk).
In fact, while the availability of consolidated components (potentially) enables the implementation of quality services to support better informed decisions, it does not clear the issues, including the regulation-related aspects, in data management, privacy, security, and surveillance. Smart city data lakes usually involve data, including georeferenced and multimedia data, about people or which allows us to track and profile people, thus potentially constitute a very interesting and profitable target for attacks: this is why the EU regulation imposes privacy-by-design and privacy-by-default prescriptions, in order to minimize risk since the beginning of the design process and with structural measures. As an obvious consequence, all architectural choices must be driven by these principles as well, that in turn will result in a change in system performance 2 : due to the need for distribution of the sensing devices among the city, the case of smart cities does exhibit a significant case for risk analysis, as the distribution makes not feasible a realistic surveillance of all the systems and devices may easily be susceptible of attacks where personal data are most exposed.
Anyway, smart cities are built over massive volumes of data and data streams, and their governance relies on correct data governance, in turn based on a sound and conscious organization of information systems, 3 with a significant role played by the data acquisition infrastructure. Privacy-by-design and privacy-by-default, as performances, should be planned and evaluated across all stages of the design and operation cycle and all subsystems of the information system, moving a big part of the decision process in the early phase and making it system-wide rather than component-wide. The most of the risk-related evaluations must be taken as soon as possible in the process, so that any aspect will be coherent with regulations and compliance can be demonstrated by documenting good practices: this implies that different architectural proposals have to be evaluated and compared far before actual executive design is available, but anticipating all possible characteristics from the executive design which may contribute significantly to risk analysis.
Consolidated practices from performance evaluation support such an early evaluation of alternative designs in terms of quantitative non-functional specifications 4 , 5 : for privacy (and security) in complex information systems, such as smart cities information systems, such an early evaluation relies on experts and a panel of experts, which may suffer of biases. In this work, we propose an evaluation framework for an early qualitative comparison of non-functional, privacy- and security-oriented constraints in architectural proposals aiming at limiting risk. The framework is defined to guide designers and assessors in comparatively ranking, with reference to different criteria, alternative designs. The proposed framework, a partial version of which has been proposed in Iacono et al., 6 consists of a first stage, in which decisions are supported by AHP and founded on weights obtained from the contributions of domain experts and application experts, and a second stage, in which the outcomes of the first stage are validated by means of a quantitative analysis based on Technique for Order of Preference by Similarity to Ideal Solution (TOPSIS) with entropy weights to verify the possible presence of biases in the first stage. The framework is demonstrated on a case study from the domain of smart cities from the PRIN PNRR 2022 project PaB-PIF, and involves the evaluation of a traditional Cloud/Edge proposal and a Blockchain-based proposal presented in detail in Campanile et al. 7
The paper is organized as follows: in the background section, the concept of smart city is explored, as well as issues about smart cities’ data and related typology. Furthermore, the methods used in the paper, namely AHP and TOPSIS, have been detailed. Subsequently, the proposed framework and the approach used to perform AHP and TOPSIS calculations have been presented. In the last section, the results have been discussed, and future work has been outlined.
Background
“The effective integration of physical, digital and human systems in the built environment to deliver sustainable, prosperous and inclusive future for its citizens” is the definition of smart cities adopted by British Standards Institute. 8 The classified challenges of smart cities that are involved in different research areas are the following: (i) privacy and security of mobile devices and services, (ii) smart city infrastructure, (iii) smart power systems and smart healthcare, (iv) framework, (v) algorithms and protocols to improve security and privacy, (vi) operational threats for smart cities, (vii) use and adoption of smart services by citizens, and (viii) use of Blockchain system within smart cities. 9 This literature review reveals also a lack in literature on privacy and security risks, as well as on the use of Blockchain technologies, areas in which key factors such as regulations are an integral part of the system and play a fundamental role in its functioning because they must generate trust among citizens who are invited to embrace this new standard of interaction and the infrastructure transition process, and this is a consequence that allows the system to work. Let us imagine people who live in a smart city and this city presents all the necessary infrastructure to collect data on the condition of vehicles and drivers. This continuous monitoring increases public safety and provides early warning of any potential anomalies that could increase the risk of road accidents. If the data is not sufficiently protected, we face a high security risk because it would be very easy for members of organized crime who target police, officials, judges, politicians, and tycoons to hack the system and breach security. 7 Satoshi Nakamoto 10 first outlined Blockchain technology as the foundation for Bitcoin, a digital currency. It uses a proof-of-work consensus technique to validate transactions on a peer-to-peer network, removing the need for a central intermediary. 11 The consensus algorithm is at the heart of Blockchain technology, ensuring that all network members agree on the ledger’s current state.12,13 Essentially, Blockchain is a digital ledger with unique qualities such as transparency, trustworthiness, immutability, traceability, decentralization, and tamper-proof nature that make it highly resistant to unauthorized changes. Because of these characteristics, several applications have been proposed in both business and academia in recent years, which also point to the secure characteristics of such technology. 14 Returning to our example, it is therefore necessary to guarantee data security and privacy so that a transition strategy can be developed and managed that ensures the integration of all components of the smart city and security: Blockchain. Security and privacy issues are often linked to governance based on inadequate risk assessment for operations. Finally, it should also be borne in mind that another key factor is the speed at which the transition takes place. Consider, for one last time, being in a smart city where sensors are scattered throughout, capable of detecting the movements of its inhabitants, and these networked devices can communicate information such as location, home address, and information about the activities they perform, perhaps to users who interact with the device’s applications. A system capable of protecting privacy must be constantly aligned with security requirements to ensure trust and well-being in smart cities. For this reason, in this highly complex system, where internet of things (IoT) is at the heart of smart cities and communication between all connected devices facilitates a high volume of sensitive data exchange, there are still unresolved challenges, and it is precisely these two pillars that research must focus on: data management and the impact on security and privacy. So, weaknesses cannot be afforded in these processes, which are subject to careful risk analysis that generally involves qualitative steps. Many approaches have been proposed in order to tackle such a situation, including quantitative methods such as the analytic hierarchy process (AHP) 15 or the Technique of Order Preference Similarity to the Ideal Solution (TOPSIS [for an in-depth view over TOPSIS, refer to Behzadian et al. 16 ]) 17 : in Chang et al., 18 an original design manufacture risk assessment has been proposed for companies data collections, Kinjo et al. 19 used the AHP to identify risky network components within the IoT components of an healthcare system, Lin et al. 20 found a set of countermeasure to protect personally identifiable information conducting an AHP-based risk assessment, showing again the elasticity of the model to adapt at very different applications. Finally, it has to be noted that several researchers use two or more different multicriteria techniques in a mixed approach or to cross-validate the results of the research.21–23
AHP method
Although AHP was first introduced in the 1980s, it is still a widely accepted and consolidated methodology to approach decision-making problems, including risk assessment in various fields. 24 The aim of the present study is to compare risks associated with Cloud and Blockchain architectures within the context of smart city data management, where AHP stands out as the methodological backbone. AHP allows us to break down the initial and complex decision problem into a hierarchical structure composed of a (i) decision goal; (ii) set of criteria (and sub-criteria); and (iii) set of alternatives (i.e. Cloud vs. Blockchain).
By means of pairwise comparisons, the idea is to compare each risk dimension, where the relative importance of one criterion over another is expressed numerically with the following Saaty’s scale.
Saaty’s scale
A key element in AHP is Saaty’s scale of absolute numbers, running from 1 to 9: 1: Equal importance 3: Moderate importance of one factor over another 5: Strong importance 7: Very strong or demonstrated importance 9: Extreme importance
Intermediate values (2, 4, 6, 8) can be used if a factor is slightly more important than the other.
At the foundation of the AHP, there is the construction of the pairwise comparison matrix
In order to check the trustworthiness of the judgments expressed with the pairwise comparisons, the consistency check is performed through the computation of the consistency index (CI):
Nevertheless, the whole AHP can be resource-intensive, particularly when the number of criteria and sub-criteria increases, as it requires up to
AHP-Express
The AHP-Express variant, proposed by Leal,
25
still lies on the basic principles of AHP, but it drastically reduces the cognitive and operational load due to pairwise comparisons. In fact, all the pairwise comparisons are not needed anymore by setting a Reference factor i is the index corresponding to the reference factor R;
Once the vector of priorities is obtained, for all factors at that level (including R), they are normalized so they sum to 1. The same procedure applies at each level of the hierarchy—criteria, sub-criteria, and so on. This resulted in a significant reduction in the total comparisons needed while preserving the essential logic of AHP.
The efficiency of AHP-Express is particularly relevant in our application, especially in the early design stages of smart city architectures, where timely and robust risk assessments are essential to perform informed decisions. The hierarchical structure ensures that the assessment captures both technical drawbacks and operational risks, allowing for a comprehensive understanding of the tradeoffs involved in selecting the most privacy-preserving and secure data management architecture for smart city applications.
Unlike the traditional AHP, which requires the CR and CI to be computed to assess the consistency of the pairwise comparison matrix, the AHP-Express variation ensures automatic consistency by design. 25 In fact, all criteria are evaluated against a single reference element; the comparison matrix that is generated is always consistent, avoiding the need for an explicit CR check. This characteristic is a major benefit in early-stage assessments since it reduces the cognitive load on specialists and avoids any inconsistencies while preserving methodological soundness.
TOPSIS
The TOPSIS technique is based on the identification of the “ideal” and “anti-ideal” solutions and calculating the distances of each alternative from these points.
26
Similarly, as in AHP, TOPSIS also uses a decision matrix with
Furthermore, the TOPSIS methodology can be decomposed into the following steps.
Normalization of the decision matrix
This step is necessary to render the criteria of various types dimensionless for comparison purposes. In the normalized matrix, the performance value
Calculation of the weighted normalized decision matrix
In TOPSIS, weights may be either heuristic—such as those supplied by a group of domain expert, which may result in biased evaluations due to their subjective nature. As an alternative, the weights may be derived from data using the entropy heuristic,
27
a method that has seen extensive application; the processes for calculating entropy are delineated as follows: Calculate entropy
Calculate the degree of diversion
Entropy weights:
At this point, it is possible to define the normalized weighted matrix as:
Determination of the ideal and anti-ideal solutions
In this phase, the ideal and anti-ideal solutions can be computed via the following definitions:
Computation of the separation measures
In this phase, the Euclidean distances of each alternative from the ideal and anti-ideal solutions can be computed, respectively:
Calculation of the relative closeness to the ideal solution
Following, it is possible to compute the relative closeness of the ideal solution is calculated, these values are between 0 and 1, and better when they get closer to 1:
Ranking of the preference order
In the final step, the alternatives are ranked from the best to the worst.
Risk assessment framework
The exposure to privacy or security-related risk is the main object of interest of risk management. Risk management is the discipline that studies and defines what is related to the general problem of risk evaluation, including related processes, and the application of its findings to any area in which risk must be identified, analyzed, evaluated, and mitigated in processes, structures, systems, organizations, or decision processes. Risk assessment is one of the practices of risk management, and focuses on the evaluation of risk factors, according to a specific framework suitable for the target problem, generally on a comparative basis. Risk assessment is at the foundation, for example, of bank credit services, of anti-bribery procedures, of the EU AI act regulation, and, of course, of privacy or security. A risk assessment framework helps decision-making by providing a structured process to identify, in a specific field or with general guidelines, potential threats and vulnerabilities, and to explore the resulting consequences in likelihood and severity, supporting with a structured approach the evaluation process and informing the decisions about risk acceptance or investment in risk mitigation.
In the context of this work, risk-related decisions are forced by the GDPR into the earliest phase of the design process, if not in the preliminary activities. In these conditions, decisions have a fundamental impact on the most general architectural decisions even before the design parameters are accessible to designers or even defined, as they may depend on the technological consequences of such decisions. A risk assessment framework suitable for smart cities is thus absolutely needed to compare alternative architectural proposals in terms of risk exposure and mitigation possibilities about security and privacy vulnerabilities, and related existing procedures or procedures to be designed, which will then contribute, once the preliminary decisions are taken, to generate functional requirements. 1
The framework supports the identification of risk factors, but the intrinsically qualitative/comparative nature of risk management requires a prioritization of risk factors, which may also not be fully independent and which also generate mutual contrast, if a mitigation action for one factor may raise the exposure to another. Prioritization may be based on the likelihood of the actual occurrence of the risk and on the severity of its impact, requiring a multi-criteria decision-making approach based risk management process.28,29 We decided to combine the benefits of AHP and TOPSIS because they are well-established and validated techniques, but also because this allows us to combine the benefits of a qualitative/expert-based approach with the ones of a quantitative/analytic approach. These two techniques are used for mutual cross-validation of the qualitative interpretation of the results to improve the quality of the process.
The framework here defined focuses on supporting the main architectural decision between founding the data management subsystem of a smart city monitoring system on a pure Blockchain-based design rather than on a mixed Blockchain/Cloud-based design. The framework is designed to break down hierarchically the decision problem into its factors (risk dimensions), according to AHP rationale, in order to guide both experts and designers in systematically assigning the most suitable rating to all risk factors both in general terms and in the specific terms of the project, respectively.
In order to identify the needed hierarchy from the perspective of GDPR privacy-by-design, it is useful to focus on its main aspects: data confidentiality, data integrity, data availability, and compliance.
The first is related to the ability of a system to protect data from unauthorized access, that is, from intentional non-authorized subjects or from side-effect exposure of data by processes, in order to minimize the likelihood of data breaches. Data may be involved in complex utilization schemas, including multiple applications, even simultaneously, so that access control mechanisms, which enable controlled and possibly journaled access, and encryption may operate as mitigation tools. With reference to this, a Blockchain-based system has the advantage of data and data access points replication, immutability of implicit journaling for modifications and transparency to encryption, while a conventional Cloud-based solution, which also potentially supports logical replication of data access points by elasticity, depends both on application-related features, controlled by designers, and centrally related and depending contributions, out of the control of designers and totally exposed to the provider, thus requiring an explicit or implicit collaboration of both subjects to prevent breaches and apply mitigation strategies.
The second concerns the ability of the system to guarantee that data are correctly preserved by the system against accidental or malicious modifications or deletions. Data in a smart city context may have legal effects and may impact fundamental rights of involved people, so this is a crucial aspect to be considered in our framework. Data must be reliable: from this point of view, Blockchain cannot be altered or corrupted, due to its intrinsic immutability, while in a Cloud-based approach, this requires proper services and their continuous and active application and control.
The third is about the actual and constant accessibility of data when needed and legitimate. Multiple purposes and unscheduled operations may require access to data in different ways and whenever needed, on different scales and with different time constants. From this point of view, Blockchain ensures, as replicated, many access points to data and basically no point of failure, but has a high latency both for append and lookup operations, because of the chained structure, while Cloud might be more exposed to temporary failures and exhibits more access flexibility, performance scalability, and availability.
Finally, compliance is the conformance of a system to regulations, in this case to GDPR. Conformance is related to an application, to be intended in the widest meaning possible, rather to an architectural solution, as it involves its use and use policies, but some features of a technology may pose general challenges: in the case of Blockchain, its nature is in general in contrast with some of the GDPR previsions, such as the right to be forgotten, if interested data are stored in the Blockchain without proper (and complex) solutions which, in some measure, do cripple some of the benefits of the Blockchain itself; Cloud-based solutions are more flexible and pose no challenges in implementing compliant procedures.
It is important to note that the immutable nature of Blockchain poses challenges related to Article 17, and so it requires architectural solutions, for example, off-chain storage or data encryption with keys that can be deleted, in order to face this controversy. 30 On the other side, Cloud-based solutions are naturally more flexible towards the GDPR articles implementation, such as Articles 20 and 32–34.
A framework for risk analysis suitable for our purposes must consider all four of these aspects to provide correct information of the decision process. These aspects are, as needed, sufficiently general to be applied before details are available for system components. The hierarchical structure needed to apply AHP to our case will assign risk relative weights considering these four aspects and help prioritize the features of the alternative architectural solutions, leveraging them. On the other side, a comparison with the application of TOPSIS to the same process will allow us to detect biases or confirm results and strengthen the value of the decision.
Approach
Let us go deeper into the needs of the evaluation process to better clarify our modus operandi. In general, risk assessment may take in a design cycle two possible roles, corresponding to two opposite collocations in the cycle: at the beginning, in the preliminary part of the design process, case in which it will not consider implementation details but will shape the large scale decisions, or in the final part, case in which a more precise evaluation of risks and their implications, and of mitigation strategies can be obtained, but, unless serious issues emerge, it cannot impact significantly on the undertaken decisions, because of the investment already made in the process. An alternative is given by the possibility of a risk analysis related to an intermediate stage of the design process, in which it is part of the validation of the outcomes of that stage, but can only have local effects, hardly capturing system-wide risks, risk propagation, and some opportunity for mitigation outside the focus of that stage.
In our case, Article 25 of the GDPR prescribes that all systems which may be involved in the processing of privacy-related data must follow the privacy-by-design principle, which, as seen, forces the design process to focus on privacy issues since its earliest phase. Although this may force the needed risk analysis to be performed when a few, and basically conceptual, details on the project are available, the risk of exposing the project and its products to the legal consequences of an accusation of a violation of the GDPR is too severe to be accepted. (The same can be said about the compliance with the privacy-by-default principle, which is anyway not the focus of this paper.) This implies working on the main architectural decisions, and, by the way, does not exclude in any way the possibility of performing a second risk analysis in the final part of the design process or locally to any critical stage. Consequently, in this approach, the analysis is performed in the first part of the design process and by using a consolidated and widely accepted analysis tool, AHP, also in the perspective of making the approach strong against possible legal accusations by exploiting, through a consistent tool, expertise consolidated in a structured evaluation process in support to a equally structured specification evaluation for the target application, bringing the only space for accusation as much as possible in the limits of professional expertise and autonomy. AHP has the advantage of condensing the power of multiple expert informed opinions in a relatively simple guide to interpret results and take decisions, providing a natural first defense line that can be used in a tribunal, as it is supported by a guarantee of integration of both qualitative and quantitative factors in the risk assessment process, which then we further improve by the “second opinion” based on TOPSIS.
Factors for the analysis
We start the definition of our hierarchy for AHP from three primary factors: the Human Factor, Internal infrastructure, and the Public infrastructure. Each factor is then further categorized into subfactors, as reported in Table 1.
AHP evaluation hierarchy.
AHP evaluation hierarchy.
AHP: analytic hierarchy process; DPV: Data Protection Violation; DMV: Data Modification Violations; HE: Human Errors; EF: Equipment Failure; SB: Security Breaches; BMV: Base Software & Middleware Vulnerabilities; PF: Power Failures; NM: Network Malfunction.
The Human Factor includes all kinds of risks which are a potential consequence of the actions of humans involved in the operations of the system: this includes unwanted actions, such as errors, and intentional violations of the system or its policies by users and operators, such as misappropriation of personal data or non-legitimate modification of data. This factor includes subfactors Data Protection Violations (DPV), Data Modification Violations (DMV), and Human Errors (HE).
The Internal Infrastructure factor includes all kinds of risks which are resulting from exposures due to the HW/SW platforms and infrastructures used for the application, including procedures and resources, which are under the authority of the legal or technical entity managing the application: for example, this includes, in our case, the Blockchain or self-managed Cloud infrastructures, or the sensors used to collect data, or the software layers implementing user or administration features. This factor includes subfactors Equipment Failure (EF) and Security Breaches (SB).
The Public Infrastructure factor includes all kinds of risks which are the result of the use of public infrastructures for computing, communications or other purposes: for example, this includes, in our case, used Cloud services, network links, power provisioning, database software or middleware, which are not under the control of the legal or technical entity managing the application. This factor includes subfactors Power Failures (PF), Network Malfunctions (NM), and Base Software & Middleware Vulnerabilities (BMV).
The overall resulting hierarchy is in Figure 1. For each of the subfactors, a comparison-based evaluation is requested for the Cloud-based and the Blockchain-based reference solution, in order to prioritize each of the solutions according to each of the subfactors and obtain an informed decision supporting GDPR compliance, which can be documented and motivated. This approach allows us to invest since the earliest phases of the development of the project, all resources towards the less risky solution, reducing connected costs.

AHP hierarchy used for early privacy risk assessment. AHP: analytic hierarchy process: DPV: Data Protection Violation; DMV: Data Modification Violations; HE: Human Errors; EF: Equipment Failure; SB: Security Breaches; BMV: Base Software & Middleware Vulnerabilities; PF: Power Failures; NM: Network Malfunction.
The entire process of the suggested risk assessment methodology is shown in Figure 2. At the beginning of the process, there is defining the system specifications and identifying reference architectural models (i.e., Cloud vs. Blockchain). Risk coefficients are then assessed for every alternative after subsystem functions are broken down into risk-related factors and sub-factors. The consequent step involves domain and privacy experts, where they are interviewed in a structured manner to extract priorities. At this point, the alternative values are calculated and an initial ranking of the candidate architectures is obtained using the AHP-Express method (red). The TOPSIS method (green) is then used to validate the results in order to increase their robustness and reduce any potential biases in expert judgments. This two-step evaluation (AHP in conjunction with TOPSIS validation) ensures both transparency and methodological rigor in the early-stage comparison of architectural alternatives, thus offering decision-makers a structured and GDPR-compliant support tool.

Workflow for the risk assessment methodology comparing Blockchain and Cloud architectures, including AHP (red) and TOPSIS validation (green). AHP: analytic hierarchy process; TOPSIS: Technique for Order of Preference by Similarity to Ideal Solution.
It is to be noted that the AHP hierarchy defined below is general enough to be suitable to be use in other medium-large-scale distributed and heterogeneous systems, also because it is meant to be applied in the very early phase of the design cycle. The expert group, in this case, has to score the performance of all alternatives in function of factors and sub-factors and to set the usual preferences between sub-factors.
Part of the work here presented, has been the development of a general application tool based on the simplified version of AHP, which allowed us to compare the two architectural paradigms under evaluation, for example, Cloud and Blockchain, considering key risk factors related to privacy, security, and regulatory compliance, while reducing the cognitive burden (or bias) induced by the standard AHP methodology, and to do so facilitate a data-driven risk assessment approach.
The tool, completely developed in Python, stands as a web-based application in Streamlit, offering user-friendly usage leveraged by a user interface that offers the possibility to input criteria weights according to Saaty’s scale. It allows for pairwise comparisons and provides the final priority vector for the alternatives’ ranking.
The tool supports the following functionalities: it implements the AHP-Express that requires only it supports the definition of the hierarchical model: users can define the objective, specify the alternatives (i.e., Blockchain vs. Cloud), and structure into hierarchy risk factors and sub-factors; it supports AHP-Express fashion pairwise comparisons: each element is compared only with a reference element within its category, reducing the number of required comparisons; it performs weight normalization and prioritization: input weights are automatically normalized to sum to 1, and priority values are calculated according to Saaty’s scale; it provides graphical representations and insights: it is possible to visualize the decision hierarchy using Graphviz and generate bar charts for factor weights and alternative scores using Matplotlib; it computes the final risk assessment and returns the ranking of the alternatives, it computes the final scores of the alternatives in comparison using a weighted aggregation method. Moreover, it also offers the possibility to download the results as CSV.
The full Python implementation of this tool is available on GitHub (https://github.com/christianriccio/AHP-risk). To test the online interactive Streamlit application, it can instead be accessed at the following link: https://ahp-risk.streamlit.app/.
Case study
To validate the approach, it has been applied to a real case, which is currently studied within the PaB-PIF PRIN PNRR 2022 research project and is here proposed in a simplified version to fit the scope of this paper, without loss of generality. Full specifications are available for interested readers in Campanile et al.7,31 The architecture that is proposed is based on a system devoted to act as a trusted, shared and privacy-aware data repository on a private Blockchain to both relief the efforts needed by participants to keep data and to constitute a platform in which data are non-repudiable and verifiable in autonomy by each participant while keeping data confidentiality, as participants are in a position of potential competing interest. On the other side, the Cloud-based alternative in which all data is stored separately by each participant and all data modifications occur by means of traditional tools and kept coherent with traditional database-oriented techniques. The purpose of the system is to keep track of all events generated from smart city devices, that is, smart roads, vehicles equipped with black boxes which interact with the system by providing their state over time, vehicle manufacturers, garages performing maintenance, insurance companies, fleet managers, and authorities. Participants of the system consequently include city and relevant national authorities, smart road concessionaires, the aforementioned vehicle manufacturers, fleet managers and insurance companies, and authorized third parties which sell added value services (e.g., to garage owners, acting as proxy in the infrastructure for them). The former architectural model runs a node of the Blockchain to access data and store new data; in the second, each of them interacts with the others and with the system by means of a Cloud-based infrastructure.
In both cases, the main components of the system are internal infrastructure for each participant, a shared public infrastructure to access the system and interact between partners, and, of course, the human interaction interface.
In the first case, the public infrastructure is used to access and manage data. In the second case, human-based interventions happen via the human interaction interface, and data management happens either entirely or in part through the internal infrastructure. A Blockchain ensures that updates are fully tracked and deletions require some ad hoc mechanism to get the same legal effect of a deletion according to the GDPR, whereas a conventional storage mechanism offers full manipulation of data, possibly with some journaling or logging stored in the system with an analogous technology. In the first scenario, data can be verified in blocks for a coherence check because the Blockchain makes it immutable, even if it is no longer logically accessible. In the second scenario, data can be altered, including through destructive overwriting and deletion operations, but the journal or log must be used to track the events over time. Before defining architectural details, a choice between these two basic design directions must be made in order to comply with the privacy-by-design and privacy-by-default principles. Because it must be made early in the design process and requires a qualitative non-absolute risk analysis, the decision is based on a comparison between these two general models, as can be seen. Additionally, it calls for a number of comparisons between the two preliminary architectural models, resulting in a long list of conflicts between them that are resolved by the expertise and experience of the decision makers. Those comparisons are guided by the hierarchical structure presented in Figure 1.
Risk assessment scores
A team of cybersecurity and privacy specialists has been assembled to carry out the qualitative risk assessment process, and the scores of each sub-factor have been compared for the two options.
First, for two reasons, the group chose not to give the first-level elements any preliminary weight. it is difficult to assign a weight to first-level factors at this stage of the project, due to the need to know the real environment details (e.g., quality of public network, energy facilities, actual hired personnel, etc.); at the end of AHP computation, the three factors will show a score (the sum of all related sub-factors), and it is possible to infer about which alternative is more sensitive to different factors, giving interesting information about the nature of privacy issues.
Because the user would have suffered the same harm in both scenarios if the associated feared events had occurred, and because the likelihood of those events occurring was the same in both options, the expert group assigned the sub-factors DPV and NM the same score.
Experts gave the alternative Blockchain a five-fold lower score for the sub-factors of DMV, EF, and SB because distributed ledgers are much more resilient to those dreaded events. This viewpoint holds that the sub-factor PF is less important for Blockchain alternatives and has been rated six times lower than the Cloud alternative.
For resilience reasons, the expert group, in the case of HE, decided that the alternative with a better (lower) score is the Blockchain-based one, which is ranked three times better than the Cloud alternative.
Lastly, because large Cloud service providers are better at maintaining updated software platforms, the Cloud alternative’s score for the sub-factor BMV has been evaluated three times higher than the Blockchain alternative ones.
Subsequently, the expert group assigned the preference weights, which have been computed individually for each sub-factor grouped in each factor.
Note that the weight values in the previous Table 2 are obtained as the normalization of the results of the AHP-Express process applied to the experts’ judgments. Those weights are referred to each sub-factor and subsequently normalized to ensure comparability across factors.
Sub-factor scores and expert weights.
Sub-factor scores and expert weights.
DPV: Data Protection Violation; DMV: Data Modification Violation; HE: Human Errors; EF: Equipment Failure; SB: Security Breaches; PF: Power Failures; NM: Network Malfunction; BMV: Base Software & Middleware Vulnerabilities.
With regard to Human Factor, the more relevant dread event is DPV, followed by HE and DMV. In the case of Internal Infrastructure factor, the sub-factor EF has been considered riskier than SB; instead, in the case of the Public Infrastructure factor, the most risky sub-factor is NM, followed by PF and BMV.
The whole framework of alternatives and weights evaluation is shown in Table 2. The final AHP calculation assessment results are shown in Table 3.
Risk scores of the two architectural alternatives.
DPV: Data Protection Violation; DMV: Data Modification Violation; HE: Human Errors; EF: Equipment Failure; SB: Security Breaches; PF: Power Failures; NM: Network Malfunction; BMV: Base Software & Middleware Vulnerabilities.
The final score has been calculated using the procedures outlined in the AHP methodology section, with the findings presented in Table 3. The findings indicate that the Cloud-based alternative is the more hazardous option, scoring 2.913, whilst the Blockchain-based alternative scores 1.994, yielding a relative ratio of around 3/2. The outcome of the risk assessment indicates that the Blockchain-based option must be considered in subsequent design phases due to its significantly lower privacy risk and cannot be dismissed based on the current findings. Therefore, it is essential to make additional assessments of both choices regarding prices and performance to determine the optimal choice for all cost/benefit considerations.
Further remarks can be deduced by observing the values of weighted first-level factors for both alternatives, for the sake of clarity shown in a bar diagram (Figure 3).

Comparison of risk level per individual first-level factor.
In Table 4, the criteria weights obtained both with the entropy method and uniform method (i.e., same weight to each criterion (1/8)) are reported. Because DPV and NM have identical scores in the decision matrix, it is expected that to have a null value of entropy, allowing to put more attention to other discriminating criteria.
Criterion weights used in the TOPSIS analysis.
Criterion weights used in the TOPSIS analysis.
TOPSIS: Technique for Order of Preference by Similarity to Ideal Solution; DPV: Data Protection Violation; DMV: Data Modification Violation; HE: Human Errors; EF: Equipment Failure; SB: Security Breaches; PF: Power Failures; NM: Network Malfunction; BMV: Base Software & Middleware Vulnerabilities.
Entropy-derived weights in Table 4 have been normalized to ensure they sum to 1, in line with the standard TOPSIS method. The resulting ranking scores and subsequent rankings are presented in Table 5.
TOPSIS ranking of the alternatives.
TOPSIS: Technique for Order of Preference by Similarity to Ideal Solution.
It is clear that the Blockchain architecture is at an advantage over the Cloud one under both weighting schemes. This supremacy is enhanced by the entropy weights.
In light of the above results, some considerations can be drawn: Sensitivity to data dispersion: the entropy scheme increases the impact of PF, DMV, and SB by down-weighting criteria that do not discriminate (DPV and NM). This is not a reflection of expert bias, but rather of actual operational differences between the two architectures. Robustness check: the ranking is invariant across weighting schemes, indicating that the superiority of Blockchain does de facto hold in any specific set of weights. Practical implications: from a design perspective, Blockchain should be retained as baseline architecture for subsequent cost-performance tradeoff studies, while mitigation actions for Cloud must address, in particular, power-failure resilience and tamper-proofing of modification operations.
Both techniques converge on the same ranking, reinforcing the credibility of the design decision to prefer Blockchain. This methodological triangulation reduces the likelihood of a ranking bias due to a single evaluation paradigm.
Table 6 aims to compare the TOPSIS outcomes with the AHP risk scores previously reported by the expert panel. AHP expresses residual risk (lower is better), whereas TOPSIS expresses proximity to the ideal solution (higher is better); nonetheless, both methods consistently place Blockchain at the top.
Blockchain versus Cloud according to AHP and TOPSIS.
AHP: analytic hierarchy process; TOPSIS: Technique for Order of Preference by Similarity to Ideal Solution.
The combined use of AHP (qualitative, expert-driven) and TOPSIS methods (quantitative, data-driven) provides a bias-resistant, cross-validated evaluation framework that is useful in the early stages of design, when detailed data about system implementation is not available. So, the proposed framework helps embed privacy-by-design and privacy-by-default principles even in the first stages of design, reducing legal and operational risks associated with non-compliance.
Future efforts will focus on integrating risk assessment with economic and performance measures (e.g., latency, throughput, and implementation cost) to facilitate comprehensive architectural decision-making. The feasibility of developing an adaptive version of the framework will also be explored to address evolving threats, regulatory changes, or system scalability over time.
Footnotes
Ethical considerations
Not applicable.
Consent to participate
Not applicable.
Consent for publication
Not applicable.
Funding
The authors disclosed receipt of the following financial support for the research, authorship, and/or publication of this article: This research was partially funded by the MUR PRIN PNRR 2022 grant number P20227W8ZC.
Declaration of conflicting interest
The authors declared no potential conflicts of interest with respect to the research, authorship, and/or publication of this article.
Data availability
Data are available on request from the authors.
