Article

Journal of Korean Society of Transportation. 31 August 2026. 703-721
https://doi.org/10.7470/jkst.2026.44.4.703

ABSTRACT


MAIN

  • INTRODUCTION

  • DATA SOURCES AND PREPROCESSING

  • METHODS

  •   1. Data Warehouse and Star Schema Data Cube

  •   2. Concept Hierarchies and OLAP Operations

  •   3. Scenario-Specific Validation Procedures

  • RESULTS AND DISCUSSION

  •   1. Case Study Context and Analytical Scenarios

  •   2. Scenario I. Organization and Driver Level Heterogeneity in Risky Driving

  •   3. Scenario II. Spatiotemporal Patterns of Risky Driving in Urban Areas

  •   4. Discussion

  • CONCLUSIONS

INTRODUCTION

Road traffic crashes remain a substantial global public health and transport safety burden, causing approximately 1.19 million deaths and 20–50 million nonfatal injuries each year and generating economic losses equivalent to roughly 3% of gross domestic product in many countries (World Health Organization, 2023). Within this broader burden, crashes involving commercial vehicles warrant particular attention because buses, taxis, and freight vehicles combine high operational intensity with large vehicle mass, thereby amplifying both injury severity and wider social and economic consequences (Pahukula et al., 2015; Nævestad et al., 2023). This concern is also evident in South Korea, where business vehicles accounted for approximately one fifth of all road crashes and road fatalities in 2024 (KoROAD Traffic Accident Analysis System (TAAS), 2026a, 2026b). These patterns underscore the need for more robust evidence-based strategies for commercial vehicle crash prevention in South Korea.

In South Korea, one of the principal institutional instruments for commercial vehicle safety management is the digital tachograph (DTG). DTG devices record second-level vehicle operation data and enable the extraction of 11 nationally standardized risky-driving event categories used for regulatory monitoring and safety education, including overspeeding, prolonged speeding, harsh acceleration, harsh braking, rapid start, abrupt stop, sharp left turn, sharp right turn, sharp U-turn, abrupt overtaking, and abrupt lane change (Korea Transportation Safety Authority, 2023). Their importance lies in their integration into the regulatory workflow. Nevertheless, current management remains largely grounded in simple event counts and standardized follow-up measures. This count-based logic is limited because accident indicators, DTG statistics, company registries, and driver registries are maintained across multiple administrative dimensions, but have not been organized into an integrated analytical framework for interpreting the 11 risky-driving categories (Ministry of Land, Infrastructure and Transport, 2025a, 2025b; Korea Transportation Safety Authority, 2026a, 2026b).

This limitation points to a broader analytical issue: event frequencies alone are insufficient to characterize the context, operating conditions, and managerial priority of risky driving. Similar event counts may indicate different safety-management problems when they are concentrated in different company types, driver groups, service areas, or time periods. Prior studies have shown that risky driving varies across organizational characteristics, driver characteristics, time of operation, and spatial context (Wang et al., 2018; Feng et al., 2016). Recent driving-behavior, DTG, and telematics studies have also shown that risky-driving patterns can be analyzed through vehicle telemetry, contextual data, clustering, hotspot identification, prediction, and explainable modeling approaches (Du et al., 2023; Zhou and Zhang, 2019; Cho et al., 2023b; Jin and Noh, 2023; Masello et al., 2023; Boylan et al., 2024; Ziakopoulos, 2024). However, much of this literature remains oriented toward event detection, prediction, or predefined modeling targets. As a result, limited practical analytical support is available for comparing and interpreting the 11 standardized DTG indicators across companies, drivers, locations, and time windows within regulatory practice.

Commercial-fleet safety management also requires company-level oversight because transportation companies are expected to monitor risky-driving behavior, driver fatigue, work pressure, and heterogeneous route environments (Knipling et al., 2003; Kim et al., 2018; Han and Lee, 2020; Han et al., 2023). However, institutional monitoring and safety-education systems often rely on standardized programs and broad company classifications, which may not fully reflect heterogeneity in company scale, fleet composition, driver experience, or spatiotemporal operating conditions (Cho et al., 2023a; Lee et al., 2020). This reinforces the need for a decision-support workflow that links standardized DTG event indicators with company, driver, spatial, temporal, and accident contexts.

The gap is therefore not a shortage of DTG indicators or machine-learning models, but the absence of an analytical workflow that (i) connects nationally standardized risky-driving indicators to company, driver, spatial, temporal, and accident information, (ii) does so at management-relevant aggregation levels, and (iii) examines whether event-burden patterns identified from DTG data are associated with accident outcomes after considering the available exposure information. A previous Korean DTG data-cube study demonstrated that the 11 standardized risky-driving indicators can be organized across company, driver, time, space, and accident dimensions using a multidimensional data structure (Jin et al., 2024). Building on that foundation, the present study shifts the focus from data-cube construction and online analytical processing (OLAP)-based description to accident linkage, exposure sensitivity, and post-OLAP statistical validation. A star schema data cube is retained as the integration structure, and staged OLAP is used as a decision-support scaffold—not the end-product—to progressively narrow management priorities across company type, fleet size, driver career structure, location, and time. Post-OLAP statistical validation then examines whether the narrowed priorities are supported in organization-level event-burden models and in exposure-adjusted accident-count models at the neighborhood×time-bin level.

This study makes three contributions. First, it extends the multidimensional DTG data-cube approach toward accident-informed safety management by integrating nationally standardized risky-driving events with accident, company, driver, spatial, and temporal attributes. In this framework, staged OLAP exploration is used to identify management-relevant event-burden patterns across company type, fleet size, driver career structure, location, and time, rather than to propose a new OLAP method. Second, it adds post-OLAP statistical validation by applying ordinary least squares (OLS) to organization-level event-burden patterns and negative-binomial generalized linear models (NB GLM) to neighborhood×time-bin accident counts with trip exposure. Third, it compares harsh acceleration, harsh braking, and speeding to examine whether standardized DTG categories show different organizational signatures and accident-linkage profiles. These contributions support differentiated, association-based commercial vehicle safety management that distinguishes event-burden hotspots from exposure-adjusted accident-count correlates.

The remainder of this paper is organized as follows. The next section describes the data sources and preprocessing procedures. The Methods section then explains the data-cube construction, concept hierarchies, OLAP workflow, and scenario-specific validation procedures. The Results and Discussion section presents two scenario-based analyses of organization-level and spatiotemporal heterogeneity, followed by the Conclusions section.

DATA SOURCES AND PREPROCESSING

The analytical dataset was created by integrating four source systems: DTG-derived risky-driving events, police-reported traffic-accident records, transportation-company information, and commercial-driver profiles. The integration procedure followed an extract-transform-load (ETL) workflow in which timestamps, spatial identifiers, company codes, and categorical attributes were standardized before the records were loaded into the warehouse. During this process, the source tables were screened for linkage completeness, anonymized identifiers were synchronized across datasets, and the variables required for the concept hierarchies were harmonized into consistent analytical categories. The case study focused on Daejeon, and the main multidimensional analyses were conducted for 2019 because that year represented the local peak in accident occurrence within the study period.

DTG-derived risky-driving events The primary dataset comprised records of 11 risky-driving event categories extracted from DTG data. In South Korea, commercial vehicles are required to be equipped with onboard units that collect second-level operational data managed by the Korea Transportation Safety Authority. The DTG records include timestamps, vehicle and company identifiers, speed, acceleration, braking signals, mileage, and GPS-based location information. The official event categories include overspeeding, prolonged speeding, harsh acceleration, harsh braking, rapid start, abrupt stop, sharp left turn, sharp right turn, sharp U-turn, abrupt overtaking, and abrupt lane change (Korea Transportation Safety Authority, 2023). In the remainder of this paper, these road-safety terms are used when referring to the corresponding official DTG event categories. Table 1 summarizes the official extraction rules used to identify the 11 risky-driving categories.

Table 1.

Official classification and extraction criteria for the 11 DTG risky-driving event categories

Risky driving behavior Extraction rules from DTG data
Overspeeding ⦁Counted when driving 20km/h over the road speed limit.
Prolonged speeding ⦁Counted when driving for more than 3 minutes and exceeding the road speed limit by 20km/h.
Harsh acceleration ⦁Bus: 6.0km/h per second above 6.0km/h.
⦁Taxi/General: 8.0km/h per second above 6.0km/h.
⦁Freight: 5.0km/h per second above 6.0km/h.
Harsh braking ⦁Bus: 9.0km/h per second over 6.0km/h.
⦁Taxi/General: 14.0km/h per second over 6.0km/h.
⦁Freight: 8.0km/h per second over 6.0km/h.
Rapid start ⦁Bus: 8.0km/h per second over 5.0km/h.
⦁Taxi/General: 10.0km/h per second under 5.0km/h.
⦁Freight: 6.0km/h per second under 5.0km/h.
Abrupt stop ⦁Bus: Decelerating 9.0km/h per second to a speed under 5.0km/h.
⦁Taxi/General: Decelerating 14.0km/h per second to a speed under 5.0km/h.
⦁Freight: Decelerating 8.0km/h per second to a speed under 5.0km/h.
Sharp left turn ⦁Bus: Cumulative direction angle 60-120° within 4 seconds at a speed over 25.0km/h.
⦁Taxi/General: Cumulative direction angle 60-120° within 3 seconds at a speed over 30.0km/h.
⦁Freight: Cumulative direction angle 60-120° within 4 seconds at a speed over 20.0km/h.
Sharp right turn ⦁Bus: Cumulative direction angle 60-120° within 4 seconds at a speed over 25.0km/h.
⦁Taxi/General: Cumulative direction angle 60-120° within 3 seconds at a speed over 30.0km/h.
⦁Freight: Cumulative direction angle 60-120° within 4 seconds at a speed over 20.0km/h.
Sharp U-turn ⦁Bus: Cumulative direction angle 160-180° within 8 seconds at a speed over 20.0km/h.
⦁Taxi/General: Cumulative direction angle 160-180° within 6 seconds at a speed over 25.0km/h.
⦁Freight: Cumulative direction angle 160-180° within 8 seconds at a speed over 15.0km/h.
Abrupt overtaking ⦁Bus: Changing direction angle within 8°/sec and cumulative angle within ±2°/sec at a speed over 30.0km/h while accelerating 3.0km/h per second.
⦁Taxi/General: Changing direction angle within 10°/sec and cumulative angle within ±2°/sec at a speed over 30.0km/h while accelerating 3.0km/h per second.
⦁Freight: Changing direction angle within 6°/sec and cumulative angle within ±2°/sec at a speed over 30.0km/h while accelerating 3.0km/h per second.
Abrupt lane change ⦁Bus: Changing direction angle within 8°/sec and cumulative angle within ±2°/sec for 5 seconds at a speed over 30.0km/h while accelerating 2.0km/h per second.
⦁Taxi/General: Changing direction angle within 10°/sec and cumulative angle within ±2°/sec for 5 seconds at a speed over 30.0km/h while accelerating 2.0km/h per second.
⦁Freight: Changing direction angle within 6°/sec and cumulative angle within ±2°/sec for 5 seconds at a speed over 30.0km/h while accelerating 2.0km/h per second.

Traffic-accident records To complement the event-based analysis, police-reported road traffic accident records provided by the Korean National Police Agency were incorporated (Korea National Police Agency, 2023). The accident records included accident time, location, severity, involved vehicle types, road and weather conditions, and driver-related attributes required for linkage and aggregation. By aligning accident time and location with DTG-derived risky-driving events, the integrated warehouse supported exploratory comparison between the distribution of risky-driving events and accident outcomes.

Transportation-company information and driver profiles Company and driver data were obtained from the Korea Transportation Safety Authority and related administrative sources. Company-level attributes included fleet size, operational type, and related identifiers required for warehouse integration. Driver-level attributes included demographic characteristics, total commercial-driving career, company tenure, and administrative safety-history variables. These datasets were transformed into categorical levels suitable for concept hierarchies, such as company-scale groups and age- or experience-based driver groupings. Because these variables served mainly as linkage keys and hierarchy-defining attributes, the detailed field dictionary is not reproduced in the main text.

Integration for multidimensional analysis The integrated warehouse linked risky-driving events to companies, driver groups, times, locations, and accident contexts. Company codes associated DTG records with transportation-company information, while harmonized temporal and spatial attributes supported matching with accident records and administrative-area hierarchies. This structure allowed OLAP queries to reflect not only event frequency but also the surrounding operational and safety context.

Analytic cohort for the Daejeon 2019 case study After ETL and record linkage, the analytic cohort comprised approximately 80 million DTG-derived risky-driving events distributed across the 11 standardized categories, 247 commercial transportation companies spanning Local/Route Bus, Charter/Tour Bus, Taxi, and Freight, and 6,814 commercial drivers. Each company was assigned to one of five fleet-size groups (G1–G5) using regulator-defined cut-offs, and each driver was mapped to career and age strata consistent with the concept hierarchies. Accident records with complete time and location fields were linked to risky-driving events and to the company registry through vehicle identifiers, timestamps, and dong-level spatial keys; records with missing company code, missing hour, or invalid dong were excluded before aggregation. For Scenario II, these records were aggregated into dong×time-bin cells using five time bins (AM peak, PM peak, evening, night, daytime), producing the analytical grid on which the NB GLM was estimated. Extreme values in events per vehicle (EPV), accidents per vehicle, and freight-event share were examined through outlier diagnostics and winsorization-based sensitivity checks.

METHODS

The analytical architecture comprised three components: data warehousing and star-schema data cube construction, concept hierarchies and staged OLAP exploration, and scenario-specific statistical validation. The objective was to support interpretable multidimensional exploration of DTG risky-driving event burdens and to examine whether OLAP-selected patterns were consistent with accident, exposure, and organizational evidence.

1. Data Warehouse and Star Schema Data Cube

Prior to constructing the data cube, the integrated source datasets were loaded into a warehouse through ETL procedures that standardized the linkage keys, time variables, spatial units, and categorical attributes required for multidimensional analysis. The warehouse was then organized as a star schema, which separates descriptive dimensions from the fact table and thereby supports consistent aggregation across management-relevant dimensions (Han et al., 2022; Kim et al., 2022). Figure 1 presents the resulting schema.

Operationally, standardized keys for time, location, company, driver, and accident attributes were assigned to each DTG-derived event record. Event counts were stored in the fact table, while dimension tables retained the categorical attributes needed for roll-up, drill-down, slice, and dice operations.

https://cdn.apub.kr/journalsite/sites/kst/2026-044-04/N0210440411/images/kst_2026_444_703_F1.jpg
Figure 1

Star schema for constructing the data cube model

Five major dimensions were defined: time, location, company, driver, and traffic accident. The company dimension represents company scale and operational type, the driver dimension captures demographic and experience-related characteristics, the time dimension supports temporal aggregation, the location dimension links events to administrative areas, and the accident dimension provides crash-related context. This design allows the same risky-driving events to be examined from multiple management-relevant perspectives without duplicating the event records themselves.

2. Concept Hierarchies and OLAP Operations

To support analysis at multiple levels of abstraction, concept hierarchies were defined for the core dimensions. A concept hierarchy represents ordered mappings from detailed categories to progressively more general categories, thereby enabling systematic movement between granular and aggregated views. In this study, company records were grouped by company subtype, company type, and fleet-size group; time was organized as day < {month < quarter; week} < year; location moved from dong to gu and city; and driver- and accident-related attributes were grouped into age/career and crash-context levels. These hierarchies were implemented through OLAP operations, allowing the analyst to move from broad summaries to detailed operational views while preserving the underlying dimensional logic (Noh et al., 2022; Noh and Yeo, 2021).

The principal OLAP operations were roll-up, drill-down, slice, dice, and pivot. Roll-up and drill-down adjust aggregation levels, while slice, dice, and pivot define or rearrange sub-cubes for comparison across dimensions. Figure 2 illustrates their organization into the staged OLAP workflow. The workflow was used for analyst-guided decision support rather than proposed as a new analytical method.

https://cdn.apub.kr/journalsite/sites/kst/2026-044-04/N0210440411/images/kst_2026_444_703_F2.jpg
Figure 2

Conceptual workflow of staged OLAP exploration used in this study

3. Scenario-Specific Validation Procedures

To complement OLAP-based exploration, two validation procedures were aligned with the scenario objectives. For Scenario I, an OLS regression was fitted at the company level among Local/Route Bus operators, using harsh-acceleration EPV, defined as EPVc=Ncevent/Ncveh for company c, where Ncevent is the 2019 event count and Ncveh is the registered fleet size, as the dependent variable, and workforce-composition indicators as independent variables: the proportion of long-career drivers (total commercial-driving career > 11 years), average total career (expressed in years throughout the paper), the proportion of senior drivers (age ≥ 60), and accidents per vehicle. For Scenario II, an NB GLM was fitted at the dong×time-bin level, with accident counts as the dependent variable, trip exposure (log trips) as the offset, and harsh-acceleration EPV, freight-event share, urban-core indicator, and time-period indicators as covariates. The urban-core indicator was defined as seven Seo-gu dong identified through spatial roll-up (Dunsan-1/2-dong, Wolpyeong-1/2/3-dong, Gwanjeo-1/2-dong). Because vehicle-kilometers-traveled (VKT) are not consistently available at the company level, EPV in Scenario I normalizes by fleet size rather than by distance-based exposure; this limitation is revisited in the Discussion.

A layered diagnostic protocol was adopted to assess distributional and inferential assumptions. For the OLS specification, multicollinearity was screened using variance inflation factors (VIF; retention threshold VIF < 5), and heteroskedasticity was addressed through HC3 robust standard errors. Residual behavior was assessed using the Shapiro–Wilk test for normality, the Breusch–Pagan test for heteroskedasticity, the Jarque–Bera test for the joint third- and fourth-moment departure, and Cook's distance for influential observations.

For the NB GLM specification, the choice of a negative-binomial rather than Poisson model was justified by a likelihood-ratio test against a Poisson submodel and by reporting the dispersion parameter α with its 95% confidence interval. Incidence-rate ratios (IRRs) with 95% confidence intervals were reported for each covariate, and pseudo-R2 measures (McFadden, Cragg–Uhler, Nagelkerke) were computed as complementary goodness-of-fit indicators. To detect residual spatial structure, global Moran’s I was computed on Pearson residuals aggregated to dong centroids using k-nearest-neighbor (k=8, row-standardized) and inverse-distance weights, with significance assessed by 999 random permutations (seed 42). When residual autocorrelation was detected, the NB GLM was re-run with a spatial-lag covariate (row-standardized k=8 neighbor average of the log accident rate), and changes in model fit and dispersion were summarized in the results. Extreme-value influence was assessed by winsorizing the top and bottom 1% of EPV and freight-event share and comparing the results with the unwinsorized specification. Group-level non-parametric comparisons used the Kruskal–Wallis test with ε2 as the effect-size measure, followed by Dunn's post-hoc test with Bonferroni correction. Pairwise top-vs-bottom quartile contrasts were summarized using the Mann–Whitney U test with rank-biserial correlation r and Cliff’s δ as distribution-free effect sizes.

Finally, to test robustness to risky-driving category choice, harsh braking and speeding EPV were used as alternative dependent variables in OLS and as alternative risky-driving predictors in NB GLM, while accident count remained the NB GLM dependent variable. This cross-category replication distinguishes category-specific patterns (e.g., fleet-scale concentration of harsh acceleration) from patterns generalizing across risky-driving behaviors. The validation procedures therefore complement OLAP: OLS assesses whether workforce-composition patterns from OLAP-based quartile comparisons generalize across the full Local/Route Bus sample, while NB GLM tests whether OLAP-identified spatiotemporal hotspots are statistically linked to accident occurrence after trip-exposure control, distinguishing raw-volume hotspots from exposure-adjusted accident-count correlates.

RESULTS AND DISCUSSION

1. Case Study Context and Analytical Scenarios

This section introduces the 2019 Daejeon case study and two scenario-based analyses of organization-level and spatiotemporal heterogeneity in DTG risky-driving event burdens and accident linkage.

Case-study year selection Approximately 8,400 traffic accidents occurred in Daejeon in 2019, representing the peak within the 2018–2023 observation period. The annual count then declined through 2022 before increasing again in 2023. 2019 was selected as the focal year because it simultaneously captures the highest local accident burden within the observation window and is the last full pre-pandemic year—thereby avoiding the large and heterogeneous disruptions to commercial vehicle operations, travel demand, and accident exposure that characterized 2020–2022. Robustness of the main findings to alternative focal years is discussed as a limitation.

This aggregate trend provides context but cannot account for heterogeneous company characteristics, driver composition, or fine-grained spatiotemporal structure. The results therefore develop two scenario-based analyses using staged OLAP exploration. Scenario I narrows organization-level management priorities through roll-up by company type, slice to the highest event-burden type, dice by fleet-size group, drill-down to individual companies, and post-OLAP OLS validation. Scenario II examines spatiotemporal hotspots through spatial roll-up and drill-down, company-type disaggregation, temporal drill-down, and post-OLAP NB GLM validation.

2. Scenario I. Organization and Driver Level Heterogeneity in Risky Driving

Scenario I progressively narrows organization-level management priorities through four staged OLAP operations, followed by post-OLAP statistical validation.

Company-type roll-up The analysis begins by rolling up the company dimension to the company-type level and comparing 2019 annual EPV across transportation-company types.

Local/Route Bus exhibits the highest EPV among all company types, with harsh acceleration accounting for 44.0% of total EPV, followed by harsh braking (19.7%) and abrupt lane change (9.3%) (Figure 3). Taxi shows the lowest overall EPV but a relatively high speeding share (29.1%). Company type thus serves as the primary classification axis, and Local/Route Bus is selected for subsequent analysis as the highest event-burden category.

https://cdn.apub.kr/journalsite/sites/kst/2026-044-04/N0210440411/images/kst_2026_444_703_F3.jpg
Figure 3

그림 제목

Fleet-size segmentation within Local/Route Bus The cube is sliced to Local/Route Bus operators and drilled down to company-scale groups to compare harsh-acceleration EPV by fleet size.

Median harsh-acceleration EPV increases monotonically from G1 (120) to G5 (20,541), with significant between-group differences (Kruskal–Wallis p=4.874×10-9) (Figure 4). G5 also shows the greatest dispersion. The interpretation is refined from “Local/Route Bus has the highest event burden” to “large-fleet Local/Route Bus operators concentrate the greatest harsh-acceleration monitoring burden.” Because EPV is normalized by registered vehicles rather than by VKT, vehicle-hours, or service frequency, this pattern should be interpreted as fleet-level event burden rather than pure risky-driving propensity.

https://cdn.apub.kr/journalsite/sites/kst/2026-044-04/N0210440411/images/kst_2026_444_703_F4.jpg
Figure 4

Distribution of harsh-acceleration EPV across fleet-size groups within Local/Route Bus companies

Workforce comparison by EPV quartile The analysis drills down to individual Local/Route Bus companies and compares the top- and bottom-quartile harsh-acceleration EPV groups across workforce and operational characteristics.

High-EPV companies are distinguished by long-career driver concentration, not elderly driver proportion (Figure 5). The high-EPV group shows significantly higher long-career driver proportion (total commercial-driving career >11 yr) (0.565 vs. 0.352, p<0.001), average total career (12.10 vs. 8.50 yr, p=0.001), and average driver age (52.60 vs. 49.66, p=0.008). Neither the proportion of drivers aged ≥60 (p=0.661) nor the turnover rate (p=0.822) differs significantly. The accident indicator requires caution: mean accidents per vehicle is higher in the low-EPV group (33.175 vs. 0.890, p=0.015), but this pattern is dominated by a denominator-sensitive outlier. One company in the low-EPV group has one registered vehicle and 631 linked accidents, yielding accidents per vehicle of 631. Median values are much less extreme (2.000 in the low-EPV quartile and 0.285 in the high-EPV quartile). The key finding is therefore that high-EPV operators are characterized by long-career workforce concentration, while the accidents-per-vehicle contrast is unstable and sensitive to exposure and linkage structure.

https://cdn.apub.kr/journalsite/sites/kst/2026-044-04/N0210440411/images/kst_2026_444_703_F5.jpg
Figure 5

Comparison of workforce and operational characteristics between high-EPV (Top-Q) and low-EPV (Bot-Q) Local/Route Bus companies

Post-OLAP validation OLS regression confirms part of these patterns across the full Local/Route Bus sample (Table 2). The HC3-robust model explains 40.3% of EPV variance (R2=0.403, F=12.15, p=1.34×10-7). Long-career driver proportion shows a significant positive association (b=24,359, p=0.036), whereas average total career is not significant after robust correction (b=507.2, p=0.398). The proportion of drivers aged ≥60, by contrast, shows a significant negative adjusted coefficient (p=0.0005) despite being non-significant in the unadjusted top-vs-bottom quartile comparison (p=0.661). This sign reversal is interpreted as a suppression-type partial effect and should be limited to harsh-acceleration EPV. A correlation check confirms that age-≥60 and long-career proportion are moderately but not perfectly correlated (r0.50), supporting this adjusted-effect interpretation.

Table 2.

HC3-robust OLS results for harsh-acceleration EPV among Local/Route Bus companies

Variable Coef. Std. Err. z p-value 95% CI
Intercept -3,445.1 2,536.2 -1.358 0.1744 [-8,416.0, 1,525.9]
% Senior Drivers (Age ≥ 60) -27,384.5 7,894.4 -3.469 0.0005*** [-42,857.2, -11,911.7]
% Long career drivers (>11 yr) 24,359.1 11,625.2 2.095 0.0361* [1,574.1, 47,144.2]
Avg. total career (years) 507.2 600.6 0.845 0.3983 [-669.9, 1,684.4]
Accidents per vehicle -17.0 930.2 -0.018 0.9854 [-1,840.1, 1,806.1]

note: Dependent variable = harsh acceleration EPV among Local/Route Bus companies. Model summary: R2 = 0.403, adjusted R2 = 0.370, F = 12.15, p = 1.34 × 10-7, n = 77.

***p < 0.001,

**p < 0.01,

*p < 0.05.

Diagnostic checks support a cautious interpretation. VIF values remain below the conventional threshold, indicating that severe collinearity does not drive the coefficient signs. HC3 robust standard errors leave the senior-driver and long-career proportions significant, whereas average total career is not significant. Robustness checks further show that the accidents-per-vehicle term is unstable: its coefficient is essentially null in the base HC3 model, non-significant in Huber robust regression (p=0.148), and sensitive to outlier removal or winsorization choices. These diagnostics support interpreting accidents per vehicle as a denominator- and linkage-sensitive control rather than as evidence that higher DTG EPV reduces accident risk.

To test category specificity, the OLS pipeline was re-run with harsh braking and speeding as alternative dependent variables. Cross-category replication shows that the workforce-composition pattern is strongest for harsh acceleration and harsh braking, but weaker for speeding. Harsh braking shows higher model fit (R2=0.559), whereas speeding has weaker fit (R2=0.126) and a different association profile. These contrasts support category-specific safety management rather than a uniform interpretation of all DTG risky-driving indicators.

In summary, harsh-acceleration event burden is associated with company type, fleet scale, and workforce career structure. This pattern extends partly to harsh braking but not uniformly to speeding. OLAP functions as an interpretive workflow that narrows management priorities, while post-OLAP statistical validation clarifies the uncertainty and category specificity of these patterns.

3. Scenario II. Spatiotemporal Patterns of Risky Driving in Urban Areas

Harsh-acceleration EPV varies sharply by location, time of day, and company type. Scenario II identifies these structures through four staged OLAP operations, followed by post-OLAP NB GLM validation.

District-level spatial drill-down The location dimension is drilled down from the city level to the district (gu) level to compare harsh-acceleration EPV across Daejeon districts.

Seo-gu is the dominant hotspot (EPV = 77.8), followed by Yuseong-gu (25.0), Dong-gu (21.7), Jung-gu (19.6), and Daedeok-gu (13.8) (Figure 6(a)). This concentration justifies finer-grained spatial examination.

Neighborhood-level drill-down The location dimension is further drilled down to the neighborhood (dong) level to identify the dominant micro-scale harsh-acceleration hotspots.

https://cdn.apub.kr/journalsite/sites/kst/2026-044-04/N0210440411/images/kst_2026_444_703_F6.jpg
Figure 6

Spatial drill-down of harsh-acceleration EPV in Daejeon

Dunsan-dong is the single dominant micro-scale hotspot (EPV = 45.1), roughly triple the next-highest neighborhood (Figure 6(b)). Secondary hotspots include Bongmyeong-dong (15.4), Wolpyeong-dong (12.6), Gayang-dong and Doma-dong (11.5 each), and Gwanjeo-dong (11.0). The district-level dominance of Seo-gu is thus driven primarily by Dunsan-dong—a pattern obscured by district-level averages.

Company-type disaggregation The cube is sliced by company type to examine whether neighborhood hotspot patterns differ across Local/Route Bus, Charter/Tour Bus, Taxi, and Freight operations.

The spatial structure differs fundamentally by type (Figure 7). Bus and taxi hotspots concentrate in the urban core: Local/Route Bus peaks in Dunsan-dong (75.9), Charter/Tour Bus likewise (33.2), and Taxi shows a moderate concentration (13.1). In contrast, Freight shows the strongest hotspots in Gwanjeo-dong (54.8), Doma-dong (46.3), and Jijok-dong (43.9)—peripheral neighborhoods linked to logistics corridors. The city's hotspot structure is thus multi-layered, reflecting distinct operational contexts rather than a single spatial pattern.

https://cdn.apub.kr/journalsite/sites/kst/2026-044-04/N0210440411/images/kst_2026_444_703_F7.jpg
Figure 7

Harsh-acceleration EPV hotspots disaggregated by transportation-company type, highlighting contrasting spatial patterns across company types

Temporal drill-down The time dimension is drilled down to the hourly level while retaining company-type and hotspot-neighborhood disaggregation.

Temporal profiles vary qualitatively across types (Figure 8). Local/Route Bus shows commute-aligned dual peaks (8–9 a.m., 5–6 p.m.) centered on Dunsan-dong. Charter/Tour Bus shows evening and late-night expansion. Taxi shows a moderate, flat urban-core rhythm. Freight shows afternoon—evening peaks (4–7 p.m.) on peripheral logistics corridors. The same EPV measure thus exhibits qualitatively different temporal structures depending on company type.

https://cdn.apub.kr/journalsite/sites/kst/2026-044-04/N0210440411/images/kst_2026_444_703_F8.jpg
Figure 8

Hourly harsh-acceleration EPV profiles for top hotspot neighborhoods, disaggregated by company type

Post-OLAP validation shows that exposure-adjusted accident counts are associated mainly with freight-event share, not harsh-acceleration EPV itself (Table 3 and Figure 9). Although raw risky-driving volume and accident counts are strongly correlated (ρ=0.823), the NB GLM with trip-exposure offset reveals a different structure. The model used 616 dong×time-bin cells covering 144 dongs and five time bins; zero-accident cells accounted for 0.97%, and accident counts were overdispersed (mean = 13.51, variance = 386.71). The offset was log trips, defined as the log number of unique trip identifiers per cell. The NB model strongly outperformed the Poisson model (LR=141.89, p5.15×10-33, α=0.072). Freight-event share shows the strongest positive association (IRR = 7.06, 95% CI [4.19, 11.92]), whereas harsh-acceleration EPV is not significant (IRR = 1.004, p=0.692). Nighttime is also not significant (IRR = 0.953, p=0.462), and urban-core cells show an IRR below one after controlling for trip exposure, freight-event share, and time-bin effects. These estimates are exposure-adjusted associations, not causal effects; aggregation bias and unobserved route-level confounding cannot be ruled out.

Table 3.

Negative-binomial GLM results for accident counts at the dong×time-bin level

Variable Coef. Std. Err. p-value IRR 95% CI
Harsh-acceleration EPV 0.0036 0.0090 0.692 1.004 [0.986, 1.021]
Freight-event share 1.9548 0.2670 <0.001 7.063 [4.185, 11.918]
Urban core -0.3861 0.0513 <0.001 0.680 [0.615, 0.752]
AM peak -0.5451 0.0565 <0.001 0.580 [0.519, 0.648]
PM peak -0.3449 0.0525 <0.001 0.708 [0.639, 0.785]
Evening -0.4620 0.0559 <0.001 0.630 [0.565, 0.703]
Night -0.0479 0.0650 0.462 0.953 [0.839, 1.083]
Alpha 0.0724 0.0109 <0.001 - [0.051, 0.094]

note: N = 616 dong×time-bin cells. Offset = log trips. Reference categories are non-urban-core cells and daytime. No zero- or missing-trip-exposure cells were excluded. For alpha, 95% CI refers to the dispersion parameter.

https://cdn.apub.kr/journalsite/sites/kst/2026-044-04/N0210440411/images/kst_2026_444_703_F9.jpg
Figure 9

Forest plot of incidence-rate ratios (IRRs) from the main-effects negative-binomial GLM predicting accident counts at the neighborhood×time-bin level, with trip exposure as offset

Specification checks support the negative-binomial count model and quantify remaining uncertainty. McFadden’s pseudo-R2 was 0.065, and Cragg-Uhler/Nagelkerke pseudo-R2 values were approximately 0.325, indicating modest explanatory contribution beyond the offset. Covariate VIFs were low, so the freight-share, urban-core, and time-bin effects were separately identified. The freight-event share confidence interval excluded one, whereas the nighttime interval included one, confirming that nighttime was not significant in the main model.

Global Moran’s I computed on Pearson residuals aggregated to dong centroids is small but significant (I=0.075, p=0.001), indicating residual spatial autocorrelation not fully absorbed by the covariates. This suggests that unmodeled neighborhood-level attributes, such as road hierarchy, signalization density, or depot placement, remain relevant. A spatial-lag sensitivity model improved log likelihood and reduced α from 0.072 to 0.039, suggesting that the main IRR estimates do not exhaust the spatial structure in accident occurrence. The estimates remain exposure-adjusted associations, not causal effects.

To test category specificity, the NB GLM was re-run with harsh braking and speeding as alternative risky-driving predictors while holding the offset, freight-share, urban-core, and time-bin structure fixed. Category-specific NB replications show that accident linkage differs across DTG indicators. Speeding EPV shows the clearest positive association with exposure-adjusted accident counts (IRR = 1.142, p=0.009), while harsh braking shows a smaller positive association (IRR = 1.046, p=0.042). Harsh acceleration is weaker and not robustly significant in the main model. This contrast supports using DTG categories differently for event-burden monitoring and accident-oriented prevention.

In summary, risky-driving event burden in Daejeon is differentiated by company type, space, and time, while exposure-adjusted accident-count association is dominated by freight-event share and residual spatial structure rather than by harsh-acceleration EPV itself. Post-OLAP NB GLM validation clarifies the difference between volume hotspots and exposure-adjusted accident-count correlates.

4. Discussion

Methodological implications Unlike black-box predictive models (Masello et al., 2023; Boylan et al., 2024), OLAP offers analyst-guided transparency that connects directly to safety-management workflows. The distinctive feature of the present study is that OLAP was designed as a staged interpretive process—not a flat group comparison—with post-OLAP statistical validation clarifying which patterns persisted beyond the selected sub-cubes and which required exposure- or category-specific interpretation. The validation layer combines organizational OLS diagnostics with spatiotemporal NB GLM diagnostics, including overdispersion and residual spatial-autocorrelation checks. The residual Moran’s I also prevents over-interpretation of the IRR estimates as exhaustive explanations of neighborhood accident structure.

Substantive implications for fleet safety management At the organization level, fleet supervision should target not merely large operators but specifically those with high harsh-acceleration event burden and concentrated long-career driver composition. HC3-robust OLS associated long-career driver proportion with harsh-acceleration EPV (R2=0.403), whereas average total career and accidents per vehicle were not statistically significant in the main specification. The workforce-composition pattern extends partly to harsh braking, while speeding shows a weaker and different association profile. The adjusted negative coefficient on the age-≥60 proportion, which reverses its unadjusted sign, is best interpreted as a suppression-type partial effect for harsh-acceleration EPV rather than as evidence that older drivers are safer in absolute terms. The policy implication is that intervention strategies should include career-stage-specific feedback and event-specific retraining rather than uniform training quotas or age-based screening.

At the spatiotemporal level, the city’s hotspot structure is multi-layered: bus and taxi hotspots concentrate in the urban core, whereas freight hotspots are associated with peripheral logistics corridors. The NB GLM shows that exposure-adjusted accident counts are most clearly associated with freight-event share, whereas harsh-acceleration EPV and nighttime are not statistically significant in the main model. Cross-category replication further shows that speeding has the clearest positive IRR linkage to accident counts once exposure is controlled. Congestion-related monitoring in urban cores remains relevant for reducing event volume, while accident-prevention priorities should give greater attention to freight-heavy and speeding-exposed contexts.

Limitations and future work Five limitations should be noted. First, the analysis covers a single city and a single focal year (Daejeon, 2019), selected because it is both the local peak-accident year and the last full pre-pandemic reference year for Korean commercial vehicle safety policy; multi-city and multi-year replication is needed to establish transferability beyond this setting. Second, EPV at the company level normalizes by fleet size rather than by vehicle-kilometers traveled, so behavioral burden and operational intensity are not fully disentangled at the organization level; a VKT- or trip-distance-based exposure denominator should be incorporated once the underlying data become consistently available. Third, the OLS and NB GLM results are associative rather than causal, and the significant residual Moran’s I indicates that unmodeled neighborhood structure (e.g., road hierarchy, signalization density, depot placement) remains; future work should adopt spatial-lag or conditional-autoregressive specifications and, where possible, quasi-experimental identification. Fourth, although the cross-category replication extends the OLAP logic to harsh braking and speeding, further category- and dimension-specific scenarios (e.g., idling, sharp lane change in relation to roadway type) remain open. Fifth, severity-oriented modeling was not fully developed; severity- stratified ordered, hurdle, or spatial count models would be valuable extensions.

Despite these limitations, the approach demonstrates that commercial vehicle risky-driving analysis becomes substantially more informative when standardized DTG indicators are embedded in a multidimensional structure with staged OLAP exploration and post-OLAP statistical validation enriched by explicit diagnostic reporting and cross-category replication.

CONCLUSIONS

This study developed a data-cube-based approach for multidimensional analysis of commercial vehicle risky driving, integrating DTG, accident, company, and driver data across five dimensions. Two scenario-based analyses demonstrated how staged OLAP exploration can progressively narrow management priorities, while post-OLAP statistical validation clarified which patterns persisted beyond the selected sub-cubes and which required exposure- or category-specific interpretation.

In Scenario I, Local/Route Bus exhibited the highest EPV among all company types; large-fleet operators concentrated disproportionate harsh-acceleration event burden; and long-career driver concentration distinguished high-EPV companies. HC3-robust OLS partially confirmed these patterns (R2=0.403): long-career driver proportion remained positively associated with harsh-acceleration EPV, whereas average total career and accidents per vehicle were not statistically significant. Cross-category replication showed that the organization-level pattern was stronger for harsh acceleration and harsh braking than for speeding. In Scenario II, bus and taxi hotspots concentrated in the urban core (Dunsan-dong) while freight hotspots shifted to peripheral logistics corridors, with qualitatively different temporal structures across company types. The NB GLM showed that exposure-adjusted accident counts were most clearly associated with freight-event share, whereas harsh-acceleration EPV and nighttime were not statistically significant in the main model. A small but significant residual Moran’s I signals unmodeled neighborhood structure and defines a concrete direction for spatial-lag extensions.

These findings support differentiated management strategies reflecting company type, fleet scale, driver career structure, and the distinction between event-burden hotspots and exposure-adjusted accident-count correlates. Future research should extend the analysis to multiple cities and years, incorporate vehicle-kilometers-traveled exposure, adopt spatial-lag specifications to address the observed residual autocorrelation, and broaden the staged OLAP exploration across the remaining risky-driving categories.

Funding

This work is supported by the Korea Agency for Infrastructure Technology Advancement (KAIA) grant funded by the Ministry of Land, Infrastructure and Transport (Grant No. RS-2026-25520951).

References

1

Boylan J., Meyer D., Chen W. S. (2024), A Systematic Review of the Use of In-Vehicle Telematics in Monitoring Driving Behaviours, Accid. Anal. Prev., 199, 107519.

10.1016/j.aap.2024.107519
2

Cho E., Kim Y., Lee S., Oh C. (2023a), Prediction of High-Risk Bus Drivers Characterized by Aggressive Driving Behavior, J. Transp. Saf. Secur., 1-23.

3

Cho S., Kim D., Khan H., Lee C. K. (2023b), Heinrich’s Law for Traffic Incidents: Using the Digital Tachograph Data to Identify Traffic Accident Hotspots, Promet-Traffic Transp., 35(6), 829-837.

10.7307/ptt.v35i6.207
4

Du Z., Deng M., Lyu N., Wang Y. (2023), A Review of Road Safety Evaluation Methods Based on Driving Behavior, J. Traffic Transp. Eng. (Engl. Ed.), 10(5), 743-761.

10.1016/j.jtte.2023.07.005
5

Feng S., Li Z., Ci Y., Zhang G. (2016), Risk Factors Affecting Fatal Bus Accident Severity: Their Impact on Different Types of Bus Drivers, Accid. Anal. Prev., 86, 29-39.

10.1016/j.aap.2015.09.025
6

Han J., Pei J., Tong H. (2022), Data Mining: Concepts and Techniques, Morgan Kaufmann, Cambridge, MA, USA.

7

Han S., Kim H. W., Leigh J. H. (2023), Improvement of Road Safety Management Systems of Local Governments in Korea After Evaluating Related Indicators, Accid. Anal. Prev., 193, 107325.

10.1016/j.aap.2023.107325
8

Han S., Lee H. (2020), Comparison of Road Safety Management Systems of Local Governments Using Indicators, Transp. Res. Rec., 2674(12), 435-446.

10.1177/0361198120960145
9

Jin J., Kim S., Kang N., Kim H., Lee S., Noh B. (2024), Multi-dimensional Analytical for Risky Driving Behaviors of Commercial Vehicles Based on Digital TachoGraph Data Using Data Cube Model, J. Korean Soc. Transp., 42(6), Korean Society of Transportation, 627-648.

10.7470/jkst.2024.42.6.627
10

Jin Z., Noh B. (2023), From Prediction to Prevention: Leveraging Deep Learning in Traffic Accident Prediction Systems, Electronics, 12(20), 4335.

10.3390/electronics12204335
11

Kim H., Jang T. W., Kim H. R., Lee S. (2018), Evaluation for Fatigue and Accident Risk of Korean Commercial Bus Drivers, Tohoku J. Exp. Med., 246(3), 191-197.

10.1620/tjem.246.191
12

Kim Y., Noh B., Yeo H. (2022), Urban Structure-based Factor Analysis of Two-wheeler Accidents Using Data Cube Model, J. Korean Soc. Transp., 40(3), Korean Society of Transportation, 358-379.

10.7470/jkst.2022.40.3.358
13

Knipling R. R., Hickman J. S., Bergoffen G. (2003), Effective Commercial Truck and Bus Safety Management Techniques, Transportation Research Board, Washington, D.C., USA.

14

Korea National Police Agency (2023), Korea National Police Agency Official Website, Accessed 2026-04-03.

15

Korea Transportation Safety Authority (2023), Digital Tachograph Risky Driving Event Criteria, Korea Transportation Safety Authority, Accessed 2026-04-03.

16

Korea Transportation Safety Authority (2026a), Annual Status of Transportation Companies, Korea Transportation Safety Authority, Accessed 2026-04-03.

17

Korea Transportation Safety Authority (2026b), General Status of Transport Drivers, Korea Transportation Safety Authority, Accessed 2026-04-03.

18

KoROAD Traffic Accident Analysis System (TAAS) (2026a), Business Vehicle Traffic Accidents, Official TAAS Webpage Reporting 2020-2024 Business Vehicle Crash Statistics, KoROAD, Accessed 2026-04-03.

19

KoROAD Traffic Accident Analysis System (TAAS) (2026b), Traffic Accident Comparison: Overall Road Crashes, Official TAAS Webpage Reporting Nationwide 2024 Road Crash Statistics, KoROAD, Accessed 2026-04-03.

20

Lee J., Yeo J., Yun I., Kang S. (2020), Factors Affecting Crash Involvement of Commercial Vehicle Drivers: Evaluation of Commercial Vehicle Drivers’ Characteristics in South Korea, J. Adv. Transp., 2020(1), 5868379.

10.1155/2020/5868379
21

Masello L., Castignani G., Sheehan B., Guillen M., Murphy F. (2023), Using Contextual Data to Predict Risky Driving Events: A Novel Methodology From Explainable Artificial Intelligence, Accid. Anal. Prev., 184, 106997.

10.1016/j.aap.2023.106997
22

Ministry of Land, Infrastructure and Transport (2025a), Operation Record Information, Public Data Portal Entry for DTG-Based Operation Record Statistics, Accessed 2026-04-03.

23

Ministry of Land, Infrastructure and Transport (2025b), Traffic Safety: Accident Indicator Statistics for Business Vehicles, Public Data Portal Entry Linked to TMACS Accident Indicator Statistics, Accessed 2026-04-03.

24

Nævestad T. O., Milch V., Blom J. (2023), Traffic Safety Effects of Economic Driving in Trucking Companies, Transp. Res. Part F Traffic Psychol. Behav., 95, 322-342.

10.1016/j.trf.2023.04.011
25

Noh B., Park H., Yeo H. (2022), Analyzing Vehicle-Pedestrian Interactions: Combining Data Cube Structure and Predictive Collision Risk Estimation Model, Accid. Anal. Prev., 165, 106539.

10.1016/j.aap.2021.106539
26

Noh B., Yeo H. (2021), SafetyCube: Framework for Potential Pedestrian Risk Analysis Using Multi-Dimensional OLAP, Accid. Anal. Prev., 155, 106104.

10.1016/j.aap.2021.106104
27

Pahukula J., Hernandez S., Unnikrishnan A. (2015), A Time of Day Analysis of Crashes Involving Large Trucks in Urban Areas, Accid. Anal. Prev., 75, 155-163.

10.1016/j.aap.2014.11.021
28

Wang X., Xing Y., Luo L., Yu R. (2018), Evaluating the Effectiveness of Behavior-Based Safety Education Methods for Commercial Vehicle Drivers, Accid. Anal. Prev., 117, 114-120.

10.1016/j.aap.2018.04.008
29

World Health Organization (2023), Global Status Report on Road Safety 2023, World Health Organization, Geneva, Switzerland, Accessed 2026-04-03.

30

Zhou T., Zhang J. (2019), Analysis of Commercial Truck Drivers’ Potentially Dangerous Driving Behaviors Based on 11-Month Digital Tachograph Data and Multilevel Modeling Approach, Accid. Anal. Prev., 132, 105256.

10.1016/j.aap.2019.105256
31

Ziakopoulos A. (2024), Analysis of Harsh Braking and Harsh Acceleration Occurrence via Explainable Imbalanced Machine Learning Using High-Resolution Smartphone Telematics and Traffic Data, Accid. Anal. Prev., 207, 107743.

10.1016/j.aap.2024.107743
페이지 상단으로 이동하기