tima::run_tima()
#> + par_def_pre_lib_sop_pub dispatched
#> ✔ par_def_pre_lib_sop_pub completed [8ms, 527 B]
#> + par_def_cre_com dispatched
#> ✔ par_def_cre_com completed [1ms, 375 B]
#> + par_def_pre_lib_sop_ecm dispatched
#> ✔ par_def_pre_lib_sop_ecm completed [1ms, 492 B]
#> + par_def_fil_ann dispatched
#> ✔ par_def_fil_ann completed [1ms, 1.34 kB]
#> + par_def_wei_ann dispatched
#> ✔ par_def_wei_ann completed [1ms, 5.33 kB]
#> + par_def_ann_spe dispatched
#> ✔ par_def_ann_spe completed [0ms, 2.66 kB]
#> + par_def_pre_fea_edg dispatched
#> ✔ par_def_pre_fea_edg completed [1ms, 706 B]
#> + par_def_exp_mzt dispatched
#> ✔ par_def_exp_mzt completed [0ms, 1.76 kB]
#> + par_def_pre_lib_rt dispatched
#> ✔ par_def_pre_lib_rt completed [0ms, 2.19 kB]
#> + par_def_pre_lib_sop_big dispatched
#> ✔ par_def_pre_lib_sop_big completed [1ms, 314 B]
#> + par_def_pre_lib_sop_mer dispatched
#> ✔ par_def_pre_lib_sop_mer completed [0ms, 7.20 kB]
#> + par_def_ann_mas dispatched
#> ✔ par_def_ann_mas completed [0ms, 12.36 kB]
#> + par_def_pre_tax dispatched
#> ✔ par_def_pre_tax completed [1ms, 1.51 kB]
#> + par_def_pre_ann_mzt dispatched
#> ✔ par_def_pre_ann_mzt completed [1ms, 1.18 kB]
#> + par_def_pre_lib_sop_hmd dispatched
#> ✔ par_def_pre_lib_sop_hmd completed [0ms, 492 B]
#> + par_def_pre_lib_spe dispatched
#> ✔ par_def_pre_lib_spe completed [1ms, 1.58 kB]
#> + par_def_pre_ann_spe dispatched
#> ✔ par_def_pre_ann_spe completed [1ms, 1.35 kB]
#> + par_def_pre_fea_com dispatched
#> ✔ par_def_pre_fea_com completed [1ms, 358 B]
#> + par_def_pre_ann_sir dispatched
#> ✔ par_def_pre_ann_sir completed [1ms, 1.97 kB]
#> + par_def_pre_ann_mzm dispatched
#> ✔ par_def_pre_ann_mzm completed [1ms, 1.32 kB]
#> + par_def_pre_lib_sop_clo dispatched
#> ✔ par_def_pre_lib_sop_clo completed [1ms, 523 B]
#> + par_def_pre_fea_tab dispatched
#> ✔ par_def_pre_fea_tab completed [1ms, 857 B]
#> + par_def_pre_ann_gnp dispatched
#> ✔ par_def_pre_ann_gnp completed [1ms, 1.31 kB]
#> + yaml_paths dispatched
#> ✔ yaml_paths completed [1ms, 18.84 kB]
#> + par_def_cre_edg_spe dispatched
#> ✔ par_def_cre_edg_spe completed [1ms, 1.39 kB]
#> + par_def_pre_lib_sop_lot dispatched
#> ✔ par_def_pre_lib_sop_lot completed [1ms, 494 B]
#> + paths dispatched
#> ✔ paths completed [2ms, 3.39 kB]
#> + lib_spe_exp_env_pre_pos dispatched
#> [2026-07-26 22:22:24.581] [INFO ] > Starting: download_file [url=https://github.com/adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/spectra/exp/enveda180_pos.rds, destination=data/interim/libraries/spectra/exp/enveda180_pos.rds]
#> Downloading 6% ■■■ 18s
#> Downloading 7% ■■■ 19s
#> Downloading 24% ■■■■■■■■ 14s
#> Downloading 39% ■■■■■■■■■■■■■ 12s
#> Downloading 55% ■■■■■■■■■■■■■■■■■■ 9s
#> Downloading 70% ■■■■■■■■■■■■■■■■■■■■■■ 6s
#> Downloading 83% ■■■■■■■■■■■■■■■■■■■■■■■■■■ 3s
#> Downloading 97% ■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■ 1s
#> Downloading 100% ■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■ 0s
#> [2026-07-26 22:22:45.002] [INFO ] [OK] Completed: download_file [size_bytes=1466741248] (20.4s)
#> ✔ lib_spe_exp_env_pre_pos completed [20.5s, 1.47 GB]
#> + lib_spe_exp_gnp_pre_pos dispatched
#> [2026-07-26 22:22:45.647] [INFO ] > Starting: download_file [url=https://github.com/adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/spectra/exp/gnps_11566051_pos.rds, destination=data/interim/libraries/spectra/exp/gnps_11566051_pos.rds]
#> Downloading 34% ■■■■■■■■■■■ 3s
#> Downloading 100% ■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■ 0s
#> [2026-07-26 22:22:50.093] [INFO ] [OK] Completed: download_file [size_bytes=341237933] (4.4s)
#> ✔ lib_spe_exp_gnp_pre_pos completed [4.4s, 341.24 MB]
#> + lib_spe_exp_mb_pre_pos dispatched
#> [2026-07-26 22:22:50.327] [INFO ] > Starting: download_file [url=https://github.com/adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/spectra/exp/massbank_202510_pos.rds, destination=data/interim/libraries/spectra/exp/massbank_202510_pos.rds]
#> [2026-07-26 22:22:50.820] [INFO ] [OK] Completed: download_file [size_bytes=17559329] (492ms)
#> ✔ lib_spe_exp_mb_pre_pos completed [495ms, 17.56 MB]
#> + lib_spe_exp_mer_pre_pos dispatched
#> [2026-07-26 22:22:50.937] [INFO ] > Starting: download_file [url=https://github.com/adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/spectra/exp/merlin_16984129_pos.rds, destination=data/interim/libraries/spectra/exp/merlin_16984129_pos.rds]
#> [2026-07-26 22:22:52.908] [INFO ] [OK] Completed: download_file [size_bytes=158718768] (2s)
#> ✔ lib_spe_exp_mer_pre_pos completed [2s, 158.72 MB]
#> + lib_spe_exp_mul_pre_pos dispatched
#> [2026-07-26 22:22:53.074] [INFO ] > Starting: download_file [url=https://github.com/adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/spectra/exp/multims2_17417089_pos.rds, destination=data/interim/libraries/spectra/exp/multims2_17417089_pos.rds]
#> [2026-07-26 22:22:53.574] [INFO ] [OK] Completed: download_file [size_bytes=13752263] (499ms)
#> ✔ lib_spe_exp_mul_pre_pos completed [501ms, 13.75 MB]
#> + lib_spe_is_nor_pre_pos dispatched
#> [2026-07-26 22:22:53.687] [INFO ] > Starting: download_file [url=https://github.com/adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/spectra/exp/isdbnormansusdat_14854025_pos.rds, destination=data/interim/libraries/spectra/is/isdbnormansusdat_14854025_pos.rds]
#> [2026-07-26 22:22:54.484] [INFO ] [OK] Completed: download_file [size_bytes=47223884] (797ms)
#> ✔ lib_spe_is_nor_pre_pos completed [799ms, 47.22 MB]
#> + lib_spe_is_wik_pre_pos dispatched
#> [2026-07-26 22:22:54.607] [INFO ] > Starting: download_file [url=https://github.com/taxonomicallyinformedannotation/tima-isdb-pos/raw/main/wikidata_5607185_pos.rds, destination=data/interim/libraries/spectra/is/wikidata_5607185_pos.rds]
#> Downloading 6% ■■■ 15s
#> Downloading 9% ■■■■ 13s
#> Downloading 32% ■■■■■■■■■■■ 9s
#> Downloading 47% ■■■■■■■■■■■■■■■ 8s
#> Downloading 62% ■■■■■■■■■■■■■■■■■■■■ 6s
#> Downloading 77% ■■■■■■■■■■■■■■■■■■■■■■■■ 4s
#> Downloading 90% ■■■■■■■■■■■■■■■■■■■■■■■■■■■■ 2s
#> Downloading 100% ■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■ 0s
#> [2026-07-26 22:23:13.534] [INFO ] [OK] Completed: download_file [size_bytes=1097022129] (18.9s)
#> ✔ lib_spe_is_wik_pre_pos completed [18.9s, 1.10 GB]
#> + lib_sop_lot dispatched
#> [2026-07-26 22:23:14.039] [INFO ] Retrieving latest version from Zenodo: 10.5281/zenodo.5794106
#> [2026-07-26 22:23:14.938] [INFO ] Downloading 260413_frozen_metadata.csv.gz from https://doi.org/10.5281/zenodo.5794106
#> [2026-07-26 22:23:14.940] [INFO ] > Starting: download_file [url=https://zenodo.org/api/records/19360665/files/260413_frozen_metadata.csv.gz/content, destination=data/source/libraries/sop/lotus.csv.gz]
#> Downloading ⠙
#> Downloading ⠹
#> Downloading ⠸
#> Downloading ⠼
#> [2026-07-26 22:23:25.216] [INFO ] [OK] Completed: download_file [size_bytes=90298678] (10.3s)
#> [2026-07-26 22:23:25.218] [INFO ] Download completed: data/source/libraries/sop/lotus.csv.gz
#> ✔ lib_sop_lot completed [11.2s, 90.30 MB]
#> Downloading ⠼
+ lib_sop_pub dispatched
#> Downloading ⠼
[2026-07-26 22:23:25.361] [INFO ] > Starting: download_file [url=https://zenodo.org/records/20439802/files/PubChemLite_CCSbase_20260529.csv?download=1, destination=data/source/libraries/sop/pubchemlite.csv]
#> Downloading ⠼
#> Downloading ⠙
#> Downloading ⠹
#> Downloading ⠸
#> Downloading ⠼
#> Downloading ⠴
#> Downloading ⠦
#> Downloading ⠧
#> [2026-07-26 22:23:44.837] [INFO ] [OK] Completed: download_file [size_bytes=293613581] (19.5s)
#> ✔ lib_sop_pub completed [19.5s, 293.61 MB]
#> Downloading ⠧
+ par_pre_par2 dispatched
#> Downloading ⠧
✔ par_pre_par2 completed [0ms, 35.02 kB]
#> Downloading ⠧
+ lib_xrefs dispatched
#> Downloading ⠧
[2026-07-26 22:23:45.186] [INFO ] Fetching compound cross-references from Wikidata / QLever
#> [2026-07-26 22:23:45.187] [INFO ] > Starting: get_compounds_xrefs [(no parameters)]
#> [2026-07-26 22:23:47.022] [WARN ] QLever request failed (possibly transient upstream error). Writing empty xrefs file: compounds.tsv.gz
#> [2026-07-26 22:23:47.039] [INFO ] > Starting: export_output [file=data/interim/xrefs/compounds.tsv.gz, n_rows=0]
#> [2026-07-26 22:23:47.041] [INFO ] [OK] Completed: export_output [size_bytes=35] (2ms)
#> ✔ lib_xrefs completed [1.9s, 35 B]
#> Downloading ⠧
+ lib_sop_hmd dispatched
#> Downloading ⠧
[2026-07-26 22:23:47.182] [INFO ] > Starting: download_file [url=https://hmdb.ca/system/downloads/current/structures.zip, destination=data/source/libraries/sop/hmdb/structures.zip]
#> Downloading ⠧
#> [2026-07-26 22:23:47.282] [WARN ] file download failed (attempt 1/3), retrying in 1s: HTTP 403 Forbidden.
#> [2026-07-26 22:23:48.322] [WARN ] file download failed (attempt 2/3), retrying in 2s: HTTP 403 Forbidden.
#> [2026-07-26 22:23:50.412] [WARN ] HMDB download failed: file download failed
#> ✖ x file download failed after retries Expected: Successful operation Received:
#> HTTP 403 Forbidden. Reason: Tried 3 times with exponential backoff Fix:
#> Possible solutions: 1. Check network connection 2. Verify server/service is
#> available 3. Check authentication credentials 4. Try again later if service
#> is down 5. Increase max_attempts if transient failures are common
#> [2026-07-26 22:23:50.413] [WARN ] HMDB download failed. Creating minimal placeholder SDF file.
#> ✔ lib_sop_hmd completed [3.3s, 340 B]
#> + lib_sop_ecm dispatched
#> [2026-07-26 22:23:50.561] [INFO ] > Starting: download_file [url=https://ecmdb.ca/download/ecmdb.json.zip, destination=data/source/libraries/sop/ecmdb.json.zip]
#> [2026-07-26 22:23:50.817] [INFO ] [OK] Completed: download_file [size_bytes=1334921] (256ms)
#> ✔ lib_sop_ecm completed [259ms, 1.33 MB]
#> + lib_spe_exp_mb_pre_sop dispatched
#> [2026-07-26 22:23:50.933] [INFO ] > Starting: download_file [url=https://github.com/Adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/sop/massbank_202510_prepared.tsv.gz, destination=data/interim/libraries/sop/massbank_202510_prepared.tsv.gz]
#> [2026-07-26 22:23:51.108] [INFO ] [OK] Completed: download_file [size_bytes=158982] (175ms)
#> ✔ lib_spe_exp_mb_pre_sop completed [178ms, 158.98 kB]
#> + lib_spe_is_nor_pre_neg dispatched
#> [2026-07-26 22:23:51.218] [INFO ] > Starting: download_file [url=https://github.com/adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/spectra/exp/isdbnormansusdat_14854025_neg.rds, destination=data/interim/libraries/spectra/is/isdbnormansusdat_14854025_neg.rds]
#> [2026-07-26 22:23:51.914] [INFO ] [OK] Completed: download_file [size_bytes=34220848] (696ms)
#> ✔ lib_spe_is_nor_pre_neg completed [698ms, 34.22 MB]
#> + lib_spe_is_wik_pre_neg dispatched
#> [2026-07-26 22:23:52.037] [INFO ] > Starting: download_file [url=https://github.com/taxonomicallyinformedannotation/tima-isdb-neg/raw/main/wikidata_5607185_neg.rds, destination=data/interim/libraries/spectra/is/wikidata_5607185_neg.rds]
#> Downloading 7% ■■■ 14s
#> Downloading 36% ■■■■■■■■■■■■ 7s
#> Downloading 62% ■■■■■■■■■■■■■■■■■■■■ 4s
#> Downloading 90% ■■■■■■■■■■■■■■■■■■■■■■■■■■■■ 1s
#> Downloading 100% ■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■ 0s
#> [2026-07-26 22:24:03.654] [INFO ] [OK] Completed: download_file [size_bytes=874199749] (11.6s)
#> ✔ lib_spe_is_wik_pre_neg completed [11.6s, 874.20 MB]
#> + test_spectra_mini dispatched
#> ✔ test_spectra_mini completed [0ms, 7.77 MB]
#> + lib_spe_exp_gnp_pre_neg dispatched
#> [2026-07-26 22:24:04.188] [INFO ] > Starting: download_file [url=https://github.com/adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/spectra/exp/gnps_11566051_neg.rds, destination=data/interim/libraries/spectra/exp/gnps_11566051_neg.rds]
#> [2026-07-26 22:24:05.853] [INFO ] [OK] Completed: download_file [size_bytes=91828026] (1.7s)
#> ✔ lib_spe_exp_gnp_pre_neg completed [1.7s, 91.83 MB]
#> + lib_spe_exp_mul_pre_neg dispatched
#> [2026-07-26 22:24:06.003] [INFO ] > Starting: download_file [url=https://github.com/adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/spectra/exp/multims2_17417089_neg.rds, destination=data/interim/libraries/spectra/exp/multims2_17417089_neg.rds]
#> [2026-07-26 22:24:06.273] [INFO ] [OK] Completed: download_file [size_bytes=2326830] (269ms)
#> ✔ lib_spe_exp_mul_pre_neg completed [271ms, 2.33 MB]
#> + lib_spe_exp_mer_pre_neg dispatched
#> [2026-07-26 22:24:06.387] [INFO ] > Starting: download_file [url=https://github.com/adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/spectra/exp/merlin_16984129_neg.rds, destination=data/interim/libraries/spectra/exp/merlin_16984129_neg.rds]
#> [2026-07-26 22:24:07.401] [INFO ] [OK] Completed: download_file [size_bytes=53975862] (1s)
#> ✔ lib_spe_exp_mer_pre_neg completed [1s, 53.98 MB]
#> + lib_spe_exp_env_pre_neg dispatched
#> [2026-07-26 22:24:07.536] [INFO ] > Starting: download_file [url=https://github.com/adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/spectra/exp/enveda180_neg.rds, destination=data/interim/libraries/spectra/exp/enveda180_neg.rds]
#> Downloading 26% ■■■■■■■■■ 3s
#> Downloading 91% ■■■■■■■■■■■■■■■■■■■■■■■■■■■■ 0s
#> Downloading 100% ■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■ 0s
#> [2026-07-26 22:24:11.669] [INFO ] [OK] Completed: download_file [size_bytes=215056852] (4.1s)
#> ✔ lib_spe_exp_env_pre_neg completed [4.1s, 215.06 MB]
#> + lib_spe_is_wik_pre_sop dispatched
#> [2026-07-26 22:24:11.859] [INFO ] > Starting: download_file [url=https://github.com/taxonomicallyinformedannotation/tima-example-files/raw/main/wikidata_spectral_5607185_prepared.tsv.gz, destination=data/interim/libraries/sop/wikidata_5607185_prepared.tsv.gz]
#> [2026-07-26 22:24:12.103] [INFO ] [OK] Completed: download_file [size_bytes=15074639] (244ms)
#> ✔ lib_spe_is_wik_pre_sop completed [246ms, 15.07 MB]
#> + lib_spe_is_nor_pre_sop dispatched
#> [2026-07-26 22:24:12.224] [INFO ] > Starting: download_file [url=https://github.com/Adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/sop/isdbnormansusdat_14854025_prepared.tsv.gz, destination=data/interim/libraries/sop/isdbnormansusdat_14854025_prepared.tsv.gz]
#> [2026-07-26 22:24:12.438] [INFO ] [OK] Completed: download_file [size_bytes=1236540] (214ms)
#> ✔ lib_spe_is_nor_pre_sop completed [217ms, 1.24 MB]
#> + lib_spe_exp_mb_pre_neg dispatched
#> [2026-07-26 22:24:12.549] [INFO ] > Starting: download_file [url=https://github.com/adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/spectra/exp/massbank_202510_neg.rds, destination=data/interim/libraries/spectra/exp/massbank_202510_neg.rds]
#> [2026-07-26 22:24:12.961] [INFO ] [OK] Completed: download_file [size_bytes=5972761] (412ms)
#> ✔ lib_spe_exp_mb_pre_neg completed [414ms, 5.97 MB]
#> + lib_spe_exp_env_pre_sop dispatched
#> [2026-07-26 22:24:13.078] [INFO ] > Starting: download_file [url=https://github.com/Adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/sop/enveda180_prepared.tsv.gz, destination=data/interim/libraries/sop/enveda180_prepared.tsv.gz]
#> [2026-07-26 22:24:13.417] [INFO ] [OK] Completed: download_file [size_bytes=2876546] (339ms)
#> ✔ lib_spe_exp_env_pre_sop completed [341ms, 2.88 MB]
#> + lib_spe_exp_mul_pre_sop dispatched
#> [2026-07-26 22:24:13.526] [INFO ] > Starting: download_file [url=https://github.com/Adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/sop/multims2_17417089_prepared.tsv.gz, destination=data/interim/libraries/sop/multims2_17417089_prepared.tsv.gz]
#> [2026-07-26 22:24:13.668] [INFO ] [OK] Completed: download_file [size_bytes=49548] (142ms)
#> ✔ lib_spe_exp_mul_pre_sop completed [144ms, 49.55 kB]
#> + lib_spe_exp_mer_pre_sop dispatched
#> [2026-07-26 22:24:13.774] [INFO ] > Starting: download_file [url=https://github.com/Adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/sop/merlin_16984129_prepared.tsv.gz, destination=data/interim/libraries/sop/merlin_16984129_prepared.tsv.gz]
#> [2026-07-26 22:24:13.973] [INFO ] [OK] Completed: download_file [size_bytes=823107] (198ms)
#> ✔ lib_spe_exp_mer_pre_sop completed [201ms, 823.11 kB]
#> + lib_spe_exp_gnp_pre_sop dispatched
#> [2026-07-26 22:24:14.085] [INFO ] > Starting: download_file [url=https://github.com/Adafede/SpectRalLibRaRies/raw/main/data/interim/libraries/sop/gnps_11566051_prepared.tsv.gz, destination=data/interim/libraries/sop/gnps_11566051_prepared.tsv.gz]
#> [2026-07-26 22:24:14.225] [INFO ] [OK] Completed: download_file [size_bytes=493387] (141ms)
#> ✔ lib_spe_exp_gnp_pre_sop completed [143ms, 493.39 kB]
#> + lib_sop_hmd_fam_raw dispatched
#> [2026-07-26 22:24:14.351] [INFO ] > Starting: download_file [url=https://www.csfmetabolome.ca/system/downloads/current/csf_metabolites_structures.zip, destination=data/source/libraries/sop/csfmetabolome/structures.zip]
#> [2026-07-26 22:24:14.522] [INFO ] [OK] Completed: download_file [size_bytes=251502] (171ms)
#> [2026-07-26 22:24:14.524] [INFO ] > Starting: download_file [url=https://www.fecalmetabolome.ca/system/downloads/current/feces_metabolites_structures.zip, destination=data/source/libraries/sop/fecalmetabolome/structures.zip]
#> [2026-07-26 22:24:14.784] [INFO ] [OK] Completed: download_file [size_bytes=3201305] (260ms)
#> [2026-07-26 22:24:14.786] [INFO ] > Starting: download_file [url=https://www.salivametabolome.ca/system/downloads/current/saliva_metabolites_structures.zip, destination=data/source/libraries/sop/salivametabolome/structures.zip]
#> [2026-07-26 22:24:14.977] [INFO ] [OK] Completed: download_file [size_bytes=622845] (191ms)
#> [2026-07-26 22:24:14.979] [INFO ] > Starting: download_file [url=https://www.serummetabolome.ca/system/downloads/current/serum_metabolites_structures.zip, destination=data/source/libraries/sop/serummetabolome/structures.zip]
#> [2026-07-26 22:24:15.632] [INFO ] [OK] Completed: download_file [size_bytes=12023792] (653ms)
#> [2026-07-26 22:24:15.635] [INFO ] > Starting: download_file [url=https://www.sweatmetabolome.ca/system/downloads/current/sweat_metabolites_structures.zip, destination=data/source/libraries/sop/sweatmetabolome/structures.zip]
#> [2026-07-26 22:24:15.756] [INFO ] [OK] Completed: download_file [size_bytes=42618] (122ms)
#> [2026-07-26 22:24:15.758] [INFO ] > Starting: download_file [url=https://www.urinemetabolome.ca/system/downloads/current/urine_metabolites_structures.zip, destination=data/source/libraries/sop/urinemetabolome/structures.zip]
#> [2026-07-26 22:24:16.017] [INFO ] [OK] Completed: download_file [size_bytes=2654043] (259ms)
#> [2026-07-26 22:24:16.019] [INFO ] > Starting: download_file [url=https://mcdb.ca/system/downloads/current/milk_metabolites_structures.zip, destination=data/source/libraries/sop/mcdb/structures.zip]
#> [2026-07-26 22:24:16.105] [WARN ] file download failed (attempt 1/3), retrying in 1s: HTTP 403 Forbidden.
#> [2026-07-26 22:24:17.150] [WARN ] file download failed (attempt 2/3), retrying in 2s: HTTP 403 Forbidden.
#> [2026-07-26 22:24:19.232] [WARN ] HMDB family download failed: file download failed
#> ✖ x file download failed after retries Expected: Successful operation Received:
#> HTTP 403 Forbidden. Reason: Tried 3 times with exponential backoff Fix:
#> Possible solutions: 1. Check network connection 2. Verify server/service is
#> available 3. Check authentication credentials 4. Try again later if service
#> is down 5. Increase max_attempts if transient failures are common
#> [2026-07-26 22:24:19.233] [WARN ] HMDB download failed. Creating minimal placeholder SDF file.
#> [2026-07-26 22:24:19.237] [INFO ] > Starting: download_file [url=https://smpdb.ca/downloads/smpdb_structures.zip, destination=data/source/libraries/sop/smpdb/structures.zip]
#> [2026-07-26 22:24:19.559] [INFO ] [OK] Completed: download_file [size_bytes=23382536] (322ms)
#> [2026-07-26 22:24:19.561] [INFO ] > Starting: download_file [url=https://mimedb.org/system/downloads/2.0/mimedb.sdf.zip, destination=data/source/libraries/sop/mimedb/structures.zip]
#> [2026-07-26 22:24:19.634] [WARN ] file download failed (attempt 1/3), retrying in 1s: HTTP 403 Forbidden.
#> [2026-07-26 22:24:20.673] [WARN ] file download failed (attempt 2/3), retrying in 2s: HTTP 403 Forbidden.
#> [2026-07-26 22:24:22.760] [WARN ] HMDB family download failed: file download failed
#> ✖ x file download failed after retries Expected: Successful operation Received:
#> HTTP 403 Forbidden. Reason: Tried 3 times with exponential backoff Fix:
#> Possible solutions: 1. Check network connection 2. Verify server/service is
#> available 3. Check authentication credentials 4. Try again later if service
#> is down 5. Increase max_attempts if transient failures are common
#> [2026-07-26 22:24:22.762] [WARN ] HMDB download failed. Creating minimal placeholder SDF file.
#> [2026-07-26 22:24:22.766] [INFO ] > Starting: download_file [url=https://t3db.ca/system/downloads/current/structures.zip, destination=data/source/libraries/sop/t3db/structures.zip]
#> [2026-07-26 22:24:22.872] [WARN ] file download failed (attempt 1/3), retrying in 1s: HTTP 403 Forbidden.
#> [2026-07-26 22:24:23.963] [WARN ] file download failed (attempt 2/3), retrying in 2s: HTTP 403 Forbidden.
#> [2026-07-26 22:24:26.051] [WARN ] HMDB family download failed: file download failed
#> ✖ x file download failed after retries Expected: Successful operation Received:
#> HTTP 403 Forbidden. Reason: Tried 3 times with exponential backoff Fix:
#> Possible solutions: 1. Check network connection 2. Verify server/service is
#> available 3. Check authentication credentials 4. Try again later if service
#> is down 5. Increase max_attempts if transient failures are common
#> [2026-07-26 22:24:26.052] [WARN ] HMDB download failed. Creating minimal placeholder SDF file.
#> [2026-07-26 22:24:26.056] [INFO ] > Starting: download_file [url=https://bovinedb.ca/system/downloads/current/structures.zip, destination=data/source/libraries/sop/bovinedb/structures.zip]
#> [2026-07-26 22:24:26.701] [INFO ] [OK] Completed: download_file [size_bytes=19260214] (645ms)
#> [2026-07-26 22:24:26.703] [INFO ] > Starting: download_file [url=https://www.ymdb.ca/system/downloads/current/ymdb.sdf.zip, destination=data/source/libraries/sop/ymdb/structures.zip]
#> [2026-07-26 22:24:26.813] [INFO ] [OK] Completed: download_file [size_bytes=1200611] (110ms)
#> [2026-07-26 22:24:26.814] [INFO ] > Starting: download_file [url=https://cannabisdatabase.ca/simple/download_compound_as_sdf, destination=data/source/libraries/sop/cannabisdatabase/compounds.sdf]
#> [2026-07-26 22:24:26.916] [WARN ] file download failed (attempt 1/3), retrying in 1s: Failed to perform HTTP request.
#> Caused by error in `curl::curl_fetch_disk()`:
#> ! SSL peer certificate or SSH remote key was not OK [cannabisdatabase.ca]:
#> SSL certificate problem: certificate has expired
#> [2026-07-26 22:24:27.989] [WARN ] file download failed (attempt 2/3), retrying in 2s: Failed to perform HTTP request.
#> Caused by error in `curl::curl_fetch_disk()`:
#> ! SSL peer certificate or SSH remote key was not OK [cannabisdatabase.ca]:
#> SSL certificate problem: certificate has expired
#> [2026-07-26 22:24:30.103] [WARN ] HMDB family download failed: file download failed
#> ✖ x file download failed after retries Expected: Successful operation Received:
#> Failed to perform HTTP request. Caused by error in `curl::curl_fetch_disk()`:
#> ! SSL peer certificate or SSH remote key was not OK [cannabisdatabase.ca]:
#> SSL certificate problem: certificate has expired Reason: Tried 3 times with
#> exponential backoff Fix: Possible solutions: 1. Check network connection 2.
#> Verify server/service is available 3. Check authentication credentials 4. Try
#> again later if service is down 5. Increase max_attempts if transient failures
#> are common
#> [2026-07-26 22:24:30.104] [WARN ] HMDB download failed. Creating minimal placeholder SDF file.
#> [2026-07-26 22:24:30.107] [WARN ] Failed to create zip file, trying alternative method
#> zip warning: missing end signature--probably not a zip file (did you
#> zip warning: remember to use binary mode when you transferred it?)
#> zip warning: (if you are trying to read a damaged archive try -F)
#>
#> zip error: Zip file structure invalid (compounds.sdf)
#> ✔ lib_sop_hmd_fam_raw completed [15.8s, 62.64 MB]
#> + par_pre_par dispatched
#> ✔ par_pre_par completed [0ms, 1.69 kB]
#> + par_fin_par2 dispatched
#> ✔ par_fin_par2 completed [2ms, 4.31 kB]
#> + lib_sop_hmd_fam_pre dispatched
#> [2026-07-26 22:24:30.462] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=CSFMETABOLOME, input=data/source/libraries/sop/csfmetabolome/structures.zip, tag=csf]
#> [2026-07-26 22:24:30.574] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/csfmetabolome_prepared.tsv.gz, n_rows=445]
#> [2026-07-26 22:24:30.579] [INFO ] [OK] Completed: export_output [size_bytes=19485] (4ms)
#> [2026-07-26 22:24:30.580] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=445] (118ms)
#> [2026-07-26 22:24:30.581] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=FECALMETABOLOME, input=data/source/libraries/sop/fecalmetabolome/structures.zip, tag=fecal]
#> [2026-07-26 22:24:32.404] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/fecalmetabolome_prepared.tsv.gz, n_rows=6810]
#> [2026-07-26 22:24:32.435] [INFO ] [OK] Completed: export_output [size_bytes=237060] (31ms)
#> [2026-07-26 22:24:32.436] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=6810] (1.9s)
#> [2026-07-26 22:24:32.437] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=SALIVAMETABOLOME, input=data/source/libraries/sop/salivametabolome/structures.zip, tag=saliva]
#> [2026-07-26 22:24:32.681] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/salivametabolome_prepared.tsv.gz, n_rows=1245]
#> [2026-07-26 22:24:32.688] [INFO ] [OK] Completed: export_output [size_bytes=47303] (7ms)
#> [2026-07-26 22:24:32.689] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=1245] (252ms)
#> [2026-07-26 22:24:32.690] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=SERUMMETABOLOME, input=data/source/libraries/sop/serummetabolome/structures.zip, tag=serum]
#> [2026-07-26 22:24:41.159] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/serummetabolome_prepared.tsv.gz, n_rows=25411]
#> [2026-07-26 22:24:41.227] [INFO ] [OK] Completed: export_output [size_bytes=812712] (68ms)
#> [2026-07-26 22:24:41.229] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=25411] (8.5s)
#> [2026-07-26 22:24:41.230] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=SWEATMETABOLOME, input=data/source/libraries/sop/sweatmetabolome/structures.zip, tag=sweat]
#> [2026-07-26 22:24:41.283] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/sweatmetabolome_prepared.tsv.gz, n_rows=89]
#> [2026-07-26 22:24:41.285] [INFO ] [OK] Completed: export_output [size_bytes=4110] (2ms)
#> [2026-07-26 22:24:41.286] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=89] (56ms)
#> [2026-07-26 22:24:41.287] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=URINEMETABOLOME, input=data/source/libraries/sop/urinemetabolome/structures.zip, tag=urine]
#> [2026-07-26 22:24:42.100] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/urinemetabolome_prepared.tsv.gz, n_rows=4364]
#> [2026-07-26 22:24:42.118] [INFO ] [OK] Completed: export_output [size_bytes=209222] (18ms)
#> [2026-07-26 22:24:42.119] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=4364] (832ms)
#> [2026-07-26 22:24:42.120] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=MCDB, input=data/source/libraries/sop/mcdb/structures.zip, tag=milk]
#> [2026-07-26 22:24:42.147] [WARN ] Empty dataframe in select_sop_columns
#> [2026-07-26 22:24:42.152] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/mcdb_prepared.tsv.gz, n_rows=0]
#> [2026-07-26 22:24:42.153] [INFO ] [OK] Completed: export_output [size_bytes=256] (1ms)
#> [2026-07-26 22:24:42.155] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=0] (35ms)
#> [2026-07-26 22:24:42.156] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=SMPDB, input=data/source/libraries/sop/smpdb/structures.zip, tag=pathway]
#> [2026-07-26 22:24:55.201] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/smpdb_prepared.tsv.gz, n_rows=49817]
#> [2026-07-26 22:24:55.317] [INFO ] [OK] Completed: export_output [size_bytes=1443937] (116ms)
#> [2026-07-26 22:24:55.318] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=49817] (13.2s)
#> [2026-07-26 22:24:55.320] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=MIMEDB, input=data/source/libraries/sop/mimedb/structures.zip, tag=microbiome]
#> [2026-07-26 22:24:55.349] [WARN ] Empty dataframe in select_sop_columns
#> [2026-07-26 22:24:55.354] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/mimedb_prepared.tsv.gz, n_rows=0]
#> [2026-07-26 22:24:55.355] [INFO ] [OK] Completed: export_output [size_bytes=256] (1ms)
#> [2026-07-26 22:24:55.356] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=0] (36ms)
#> [2026-07-26 22:24:55.357] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=T3DB, input=data/source/libraries/sop/t3db/structures.zip, tag=toxin]
#> [2026-07-26 22:24:55.384] [WARN ] Empty dataframe in select_sop_columns
#> [2026-07-26 22:24:55.389] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/t3db_prepared.tsv.gz, n_rows=0]
#> [2026-07-26 22:24:55.391] [INFO ] [OK] Completed: export_output [size_bytes=256] (1ms)
#> [2026-07-26 22:24:55.392] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=0] (35ms)
#> [2026-07-26 22:24:55.393] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=BOVINEDB, input=data/source/libraries/sop/bovinedb/structures.zip, tag=NA]
#> [2026-07-26 22:25:08.884] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/bovinedb_prepared.tsv.gz, n_rows=51684]
#> [2026-07-26 22:25:09.021] [INFO ] [OK] Completed: export_output [size_bytes=1568975] (137ms)
#> [2026-07-26 22:25:09.023] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=51684] (13.6s)
#> [2026-07-26 22:25:09.024] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=YMDB, input=data/source/libraries/sop/ymdb/structures.zip, tag=NA]
#> [2026-07-26 22:25:09.488] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/ymdb_prepared.tsv.gz, n_rows=2024]
#> [2026-07-26 22:25:09.500] [INFO ] [OK] Completed: export_output [size_bytes=83615] (12ms)
#> [2026-07-26 22:25:09.501] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=2024] (477ms)
#> [2026-07-26 22:25:09.502] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=CANNABISDATABASE, input=data/source/libraries/sop/cannabisdatabase/compounds.sdf, tag=NA]
#> [2026-07-26 22:25:09.528] [WARN ] Empty dataframe in select_sop_columns
#> [2026-07-26 22:25:09.533] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/cannabisdatabase_prepared.tsv.gz, n_rows=0]
#> [2026-07-26 22:25:09.535] [INFO ] [OK] Completed: export_output [size_bytes=256] (1ms)
#> [2026-07-26 22:25:09.536] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=0] (34ms)
#> ✔ lib_sop_hmd_fam_pre completed [39.1s, 4.43 MB]
#> + par_fin_par dispatched
#> ✔ par_fin_par completed [1ms, 341 B]
#> + par_usr_cre_com dispatched
#> ✔ par_usr_cre_com completed [1.7s, 200 B]
#> + par_usr_pre_lib_sop_ecm dispatched
#> ✔ par_usr_pre_lib_sop_ecm completed [1.6s, 176 B]
#> + par_usr_fil_ann dispatched
#> ✔ par_usr_fil_ann completed [1.6s, 808 B]
#> + par_usr_pre_lib_sop_pub dispatched
#> ✔ par_usr_pre_lib_sop_pub completed [1.6s, 195 B]
#> + par_usr_pre_fea_edg dispatched
#> ✔ par_usr_pre_fea_edg completed [1.6s, 328 B]
#> + par_usr_wei_ann dispatched
#> ✔ par_usr_wei_ann completed [1.6s, 1.79 kB]
#> + par_usr_ann_spe dispatched
#> ✔ par_usr_ann_spe completed [1.6s, 1.46 kB]
#> + par_usr_pre_lib_sop_mer dispatched
#> ✔ par_usr_pre_lib_sop_mer completed [1.6s, 3.10 kB]
#> + par_usr_exp_mzt dispatched
#> ✔ par_usr_exp_mzt completed [1.6s, 425 B]
#> + par_usr_pre_lib_rt dispatched
#> ✔ par_usr_pre_lib_rt completed [1.6s, 487 B]
#> + par_usr_ann_mas dispatched
#> ✔ par_usr_ann_mas completed [1.6s, 4.72 kB]
#> + par_usr_pre_lib_sop_big dispatched
#> ✔ par_usr_pre_lib_sop_big completed [1.6s, 107 B]
#> + par_usr_pre_tax dispatched
#> ✔ par_usr_pre_tax completed [1.6s, 438 B]
#> + par_usr_pre_ann_sir dispatched
#> ✔ par_usr_pre_ann_sir completed [1.6s, 859 B]
#> + par_usr_pre_lib_spe dispatched
#> ✔ par_usr_pre_lib_spe completed [1.6s, 322 B]
#> + par_usr_pre_ann_spe dispatched
#> ✔ par_usr_pre_ann_spe completed [1.6s, 656 B]
#> + par_usr_pre_ann_mzm dispatched
#> ✔ par_usr_pre_ann_mzm completed [1.6s, 635 B]
#> + par_usr_pre_fea_com dispatched
#> ✔ par_usr_pre_fea_com completed [1.6s, 200 B]
#> + par_usr_pre_ann_mzt dispatched
#> ✔ par_usr_pre_ann_mzt completed [1.6s, 546 B]
#> + par_usr_pre_fea_tab dispatched
#> ✔ par_usr_pre_fea_tab completed [1.7s, 274 B]
#> + par_usr_pre_lib_sop_hmd dispatched
#> ✔ par_usr_pre_lib_sop_hmd completed [1.6s, 178 B]
#> + par_usr_pre_lib_sop_clo dispatched
#> ✔ par_usr_pre_lib_sop_clo completed [1.6s, 267 B]
#> + par_usr_pre_ann_gnp dispatched
#> ✔ par_usr_pre_ann_gnp completed [1.6s, 633 B]
#> + par_usr_cre_edg_spe dispatched
#> ✔ par_usr_cre_edg_spe completed [1.6s, 425 B]
#> + par_usr_pre_lib_sop_lot dispatched
#> ✔ par_usr_pre_lib_sop_lot completed [1.6s, 174 B]
#> + par_cre_com dispatched
#> ✔ par_cre_com completed [2ms, 191 B]
#> + par_pre_lib_sop_ecm dispatched
#> ✔ par_pre_lib_sop_ecm completed [1ms, 191 B]
#> + par_fil_ann dispatched
#> ✔ par_fil_ann completed [2ms, 373 B]
#> + par_pre_lib_sop_pub dispatched
#> ✔ par_pre_lib_sop_pub completed [2ms, 193 B]
#> + par_pre_fea_edg dispatched
#> ✔ par_pre_fea_edg completed [2ms, 243 B]
#> + par_wei_ann dispatched
#> ✔ par_wei_ann completed [3ms, 950 B]
#> + par_ann_spe dispatched
#> ✔ par_ann_spe completed [2ms, 583 B]
#> + par_pre_lib_sop_mer dispatched
#> ✔ par_pre_lib_sop_mer completed [2ms, 854 B]
#> + par_exp_mzt dispatched
#> ✔ par_exp_mzt completed [1ms, 270 B]
#> + par_pre_lib_rt dispatched
#> ✔ par_pre_lib_rt completed [2ms, 376 B]
#> + par_ann_mas dispatched
#> ✔ par_ann_mas completed [3ms, 1.99 kB]
#> + par_pre_lib_sop_big dispatched
#> ✔ par_pre_lib_sop_big completed [1ms, 155 B]
#> + par_pre_tax dispatched
#> ✔ par_pre_tax completed [2ms, 330 B]
#> + par_pre_ann_sir dispatched
#> ✔ par_pre_ann_sir completed [1ms, 435 B]
#> + par_pre_lib_spe dispatched
#> ✔ par_pre_lib_spe completed [2ms, 406 B]
#> + par_pre_ann_spe dispatched
#> ✔ par_pre_ann_spe completed [1ms, 323 B]
#> + par_pre_ann_mzm dispatched
#> ✔ par_pre_ann_mzm completed [2ms, 329 B]
#> + par_pre_fea_com dispatched
#> ✔ par_pre_fea_com completed [1ms, 184 B]
#> + par_pre_ann_mzt dispatched
#> ✔ par_pre_ann_mzt completed [1ms, 313 B]
#> + par_pre_fea_tab dispatched
#> ✔ par_pre_fea_tab completed [1ms, 279 B]
#> + par_pre_lib_sop_hmd dispatched
#> ✔ par_pre_lib_sop_hmd completed [1ms, 191 B]
#> + par_pre_lib_sop_clo dispatched
#> ✔ par_pre_lib_sop_clo completed [1ms, 232 B]
#> + par_pre_ann_gnp dispatched
#> ✔ par_pre_ann_gnp completed [1ms, 325 B]
#> + par_cre_edg_spe dispatched
#> ✔ par_cre_edg_spe completed [2ms, 373 B]
#> + par_pre_lib_sop_lot dispatched
#> ✔ par_pre_lib_sop_lot completed [1ms, 186 B]
#> + lib_sop_ecm_pre dispatched
#> [2026-07-26 22:25:56.940] [INFO ] Preparing ECMDB structure-organism pairs
#> [2026-07-26 22:25:57.605] [INFO ] Exporting parameters to: data/interim/params/260726_222557_prepare_libraries_sop_ecmdb.yaml
#> [2026-07-26 22:25:57.607] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/ecmdb_prepared.tsv.gz, n_rows=3760]
#> [2026-07-26 22:25:57.622] [INFO ] [OK] Completed: export_output [size_bytes=165776] (15ms)
#> ✔ lib_sop_ecm_pre completed [683ms, 165.78 kB]
#> + lib_sop_pub_pre dispatched
#> [2026-07-26 22:25:57.766] [INFO ] > Starting: prepare_libraries_sop_pubchemlite [input=data/source/libraries/sop/pubchemlite.csv]
#> [2026-07-26 22:26:10.390] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/pubchemlite_prepared.tsv.gz, n_rows=566689]
#> [2026-07-26 22:26:12.221] [INFO ] [OK] Completed: export_output [size_bytes=29385634] (1.8s)
#> [2026-07-26 22:26:12.223] [INFO ] [OK] Completed: prepare_libraries_sop_pubchemlite [n_pairs=566689] (14.5s)
#> ✔ lib_sop_pub_pre completed [14.5s, 29.39 MB]
#> + input_spectra dispatched
#> ✔ input_spectra completed [0ms, 7.77 MB]
#> + lib_sop_mer_npc_cache dispatched
#> [2026-07-26 22:26:13.005] [INFO ] > Starting: download_file [url=https://github.com/Adafede/marimo/raw/refs/heads/main/apps/public/npclassifier/npclassifier_cache.csv, destination=data/interim/libraries/sop/merged/structures/taxonomies/npc.tsv.gz]
#> Downloading 44% ■■■■■■■■■■■■■■ 1s
#> Downloading 100% ■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■■ 0s
#> [2026-07-26 22:26:15.720] [INFO ] [OK] Completed: download_file [size_bytes=201875293] (2.7s)
#> ✔ lib_sop_mer_npc_cache completed [2.7s, 201.88 MB]
#> + lib_sop_mer_cla_cache dispatched
#> [2026-07-26 22:26:15.966] [INFO ] > Starting: download_file [url=https://github.com/Adafede/marimo/raw/refs/heads/main/apps/public/classyfire/classyfire_cache.csv, destination=data/interim/libraries/sop/merged/structures/taxonomies/classyfire_cache.csv]
#> [2026-07-26 22:26:17.797] [INFO ] [OK] Completed: download_file [size_bytes=143087266] (1.8s)
#> ✔ lib_sop_mer_cla_cache completed [1.8s, 143.09 MB]
#> + lib_sop_mer_str_pro dispatched
#> [2026-07-26 22:26:18.002] [INFO ] > Starting: download_file [url=https://github.com/taxonomicallyinformedannotation/tima-example-files/raw/main/processed.csv.gz, destination=data/interim/libraries/sop/merged/structures/processed.csv.gz]
#> [2026-07-26 22:26:18.764] [INFO ] [OK] Completed: download_file [size_bytes=96419862] (762ms)
#> ✔ lib_sop_mer_str_pro completed [765ms, 96.42 MB]
#> + lib_rt dispatched
#> [2026-07-26 22:26:18.949] [INFO ] Preparing retention time libraries
#> [2026-07-26 22:26:18.961] [WARN ] No retention time library found, returning empty retention time and sop tables.
#> [2026-07-26 22:26:19.007] [INFO ] Exporting parameters to: data/interim/params/260726_222619_prepare_libraries_rt.yaml
#> [2026-07-26 22:26:19.008] [INFO ] > Starting: export_output [file=data/interim/libraries/rt/prepared.tsv.gz, n_rows=1]
#> [2026-07-26 22:26:19.010] [INFO ] [OK] Completed: export_output [size_bytes=86] (2ms)
#> [2026-07-26 22:26:19.014] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/rt_prepared.tsv.gz, n_rows=1]
#> [2026-07-26 22:26:19.015] [INFO ] [OK] Completed: export_output [size_bytes=105] (2ms)
#> ✔ lib_rt completed [69ms, 191 B]
#> + lib_sop_big_pre dispatched
#> [2026-07-26 22:26:19.159] [INFO ] Preparing BiGG structure-organism pairs
#> [2026-07-26 22:26:44.563] [INFO ] > Starting: process_smiles [n_structures=1425]
#> [2026-07-26 22:26:44.564] [INFO ] Processing SMILES with RDKit
#> Downloading uv...Done!
#> Downloading cpython-3.12.13-linux-x86_64-gnu (download) (32.6MiB)
#> Downloaded cpython-3.12.13-linux-x86_64-gnu (download)
#> Downloading numpy (15.9MiB)
#> Downloading pillow (6.6MiB)
#> Downloading rdkit (35.7MiB)
#> Downloaded pillow
#> Downloaded numpy
#> Downloaded rdkit
#> Installed 5 packages in 29ms
#> [2026-07-26 22:26:49.443] [INFO ] Processing 1424 new SMILES with RDKit
#> [2026-07-26 22:26:49.444] [INFO ] Starting SMILES processing pipeline
#> [2026-07-26 22:26:49.445] [INFO ] Input: /tmp/RtmpZ8Qf0D/file26e56308d916.smi
#> [2026-07-26 22:26:49.445] [INFO ] Output: /tmp/RtmpZ8Qf0D/file26e5557131c6.csv.gz
#> [2026-07-26 22:26:49.445] [INFO ] Input file validated: /tmp/RtmpZ8Qf0D/file26e56308d916.smi
#> [2026-07-26 22:26:49.445] [INFO ] Output file validated: /tmp/RtmpZ8Qf0D/file26e5557131c6.csv.gz
#> [2026-07-26 22:26:49.445] [INFO ] Processing parameters: workers=8, batch_size=1000, progress_interval=10000
#> [2026-07-26 22:26:49.445] [INFO ] SMILES supplier initialized
#> [2026-07-26 22:26:51.594] [INFO ] Processing complete. Total molecules processed: 1424
#> [2026-07-26 22:26:51.636] [INFO ] Successfully processed 1424 SMILES
#> [2026-07-26 22:26:51.646] [INFO ] [OK] Completed: process_smiles [n_processed=1424] (7.1s)
#> [2026-07-26 22:26:59.265] [INFO ] > Starting: process_smiles [n_structures=2085]
#> [2026-07-26 22:26:59.266] [INFO ] Processing SMILES with RDKit
#> [2026-07-26 22:26:59.277] [INFO ] Processing 1242 new SMILES with RDKit
#> [2026-07-26 22:26:59.279] [INFO ] Starting SMILES processing pipeline
#> [2026-07-26 22:26:59.279] [INFO ] Input: /tmp/RtmpZ8Qf0D/file26e5c4cb899.smi
#> [2026-07-26 22:26:59.279] [INFO ] Output: /tmp/RtmpZ8Qf0D/file26e5497a80d0.csv.gz
#> [2026-07-26 22:26:59.279] [INFO ] Input file validated: /tmp/RtmpZ8Qf0D/file26e5c4cb899.smi
#> [2026-07-26 22:26:59.279] [INFO ] Output file validated: /tmp/RtmpZ8Qf0D/file26e5497a80d0.csv.gz
#> [2026-07-26 22:26:59.279] [INFO ] Processing parameters: workers=8, batch_size=1000, progress_interval=10000
#> [2026-07-26 22:26:59.279] [INFO ] SMILES supplier initialized
#> [2026-07-26 22:27:01.159] [INFO ] Processing complete. Total molecules processed: 1242
#> [2026-07-26 22:27:01.249] [INFO ] Successfully processed 1242 SMILES
#> [2026-07-26 22:27:01.259] [INFO ] [OK] Completed: process_smiles [n_processed=1242] (2s)
#> [2026-07-26 22:27:01.357] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/bigg_prepared.tsv.gz, n_rows=2359]
#> [2026-07-26 22:27:01.371] [INFO ] [OK] Completed: export_output [size_bytes=81970] (14ms)
#> ✔ lib_sop_big_pre completed [42.2s, 81.97 kB]
#> + lib_spe_exp_int_pre dispatched
#> [2026-07-26 22:27:01.691] [INFO ] > Starting: prepare_libraries_spectra [library_name=internal, n_input_files=1]
#> [2026-07-26 22:27:01.698] [WARN ] Input file(s) not found; creating empty library template
#> [2026-07-26 22:27:03.737] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/internal_prepared.tsv.gz, n_rows=1]
#> [2026-07-26 22:27:03.739] [INFO ] [OK] Completed: export_output [size_bytes=79] (2ms)
#> [2026-07-26 22:27:03.813] [INFO ] Exporting parameters to: data/interim/params/260726_222703_prepare_libraries_spectra.yaml
#> [2026-07-26 22:27:03.814] [INFO ] [OK] Completed: prepare_libraries_spectra [n_structures=1, n_spectra_total=2, files_exported=3] (2.1s)
#> ✔ lib_spe_exp_int_pre completed [2.1s, 1.28 kB]
#> + input_features dispatched
#> ✔ input_features completed [0ms, 451.55 kB]
#> + lib_sop_hmd_pre dispatched
#> [2026-07-26 22:27:04.678] [INFO ] > Starting: prepare_libraries_sop_hmdb_like [source=HMDB, input=data/source/libraries/sop/hmdb/structures.zip, tag=NA]
#> [2026-07-26 22:27:04.706] [WARN ] Empty dataframe in select_sop_columns
#> [2026-07-26 22:27:04.711] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/hmdb_prepared.tsv.gz, n_rows=0]
#> [2026-07-26 22:27:04.713] [INFO ] [OK] Completed: export_output [size_bytes=256] (1ms)
#> [2026-07-26 22:27:04.714] [INFO ] [OK] Completed: prepare_libraries_sop_hmdb_like [n_pairs=0] (36ms)
#> ✔ lib_sop_hmd_pre completed [37ms, 256 B]
#> + lib_sop_clo_pre dispatched
#> [2026-07-26 22:27:05.143] [INFO ] Preparing closed structure-organism pairs library
#> [2026-07-26 22:27:05.145] [WARN ] Closed resource not accessible at: ~/Git/lotus-processor/data/processed/240412_closed_metadata.csv.gz. Returning empty template instead.
#> [2026-07-26 22:27:05.161] [INFO ] Exporting parameters to: data/interim/params/260726_222705_prepare_libraries_sop_closed.yaml
#> [2026-07-26 22:27:05.163] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/closed_prepared.tsv.gz, n_rows=1]
#> [2026-07-26 22:27:05.165] [INFO ] [OK] Completed: export_output [size_bytes=277] (1ms)
#> ✔ lib_sop_clo_pre completed [23ms, 277 B]
#> + lib_sop_lot_pre dispatched
#> [2026-07-26 22:27:05.579] [INFO ] > Starting: prepare_libraries_sop_lotus [input=data/source/libraries/sop/lotus.csv.gz]
#> [2026-07-26 22:27:13.103] [INFO ] [OK] Completed: prepare_libraries_sop_lotus [n_pairs=677545] (7.5s)
#> [2026-07-26 22:27:13.105] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/lotus_prepared.tsv.gz, n_rows=677545]
#> [2026-07-26 22:27:16.645] [INFO ] [OK] Completed: export_output [size_bytes=49541873] (3.5s)
#> ✔ lib_sop_lot_pre completed [11.1s, 49.54 MB]
#> + fea_edg_spe dispatched
#> [2026-07-26 22:27:17.243] [INFO ] > Starting: create_edges_spectra [method=gnps, n_input_files=1]
#> [2026-07-26 22:27:17.277] [INFO ] Creating spectral similarity network edges
#> [2026-07-26 22:27:17.279] [INFO ] Importing spectra from: data/source/example_spectra.mgf
#> [2026-07-26 22:27:17.306] [INFO ] Reading MGF file (7.41 MB) with optimized parser: data/source/example_spectra.mgf
#> [2026-07-26 22:27:19.270] [INFO ] Processed 10000 spectra...
#> [2026-07-26 22:27:20.888] [INFO ] Total spectra read: 16282
#> [2026-07-26 22:27:27.494] [INFO ] Loaded 16282 spectra from file
#> [2026-07-26 22:27:27.515] [INFO ] Combining replicate spectra by FEATURE_ID
#> [2026-07-26 22:27:30.267] [INFO ] Combined replicates: 12195 -> 4087 spectra
#> [2026-07-26 22:27:30.301] [INFO ] Sanitizing 4087 spectra (cutoff: 0)
#> [2026-07-26 22:27:31.429] [INFO ] Sanitization complete: 3660/4087 spectra retained (89.6%, 427 removed)
#> [2026-07-26 22:27:31.430] [INFO ] Import complete: 3660 spectra ready for analysis
#> [2026-07-26 22:27:31.431] [INFO ] ======================================
#> [2026-07-26 22:27:31.432] [INFO ] Take yourself a break, you deserve it.
#> [2026-07-26 22:27:31.433] [INFO ] ======================================
#> [2026-07-26 22:27:31.435] [INFO ] > Starting: create_edges [n_spectra=3660, method=gnps, threshold=NULL, min_peaks=NULL]
#> [2026-07-26 22:27:32.487] [INFO ] Processed 500 / 3659 queries
#> [2026-07-26 22:27:33.406] [INFO ] Processed 1000 / 3659 queries
#> [2026-07-26 22:27:34.068] [INFO ] Processed 1500 / 3659 queries
#> [2026-07-26 22:27:34.612] [INFO ] Processed 2000 / 3659 queries
#> [2026-07-26 22:27:35.048] [INFO ] Processed 2500 / 3659 queries
#> [2026-07-26 22:27:35.366] [INFO ] Processed 3000 / 3659 queries
#> [2026-07-26 22:27:35.570] [INFO ] Processed 3500 / 3659 queries
#> [2026-07-26 22:27:35.595] [WARN ] Sanitized 3660/3660 spectra on-demand before edge scoring.
#> [2026-07-26 22:27:35.596] [INFO ] Here is the distribution of edge similarity scores (0.1 bins):
#> [2026-07-26 22:27:35.598] [INFO ]
#> bin N Pct
#> [0,0.1] 4838216 72.26%
#> (0.1,0.2] 1077987 16.10%
#> (0.2,0.3] 418929 6.26%
#> (0.3,0.4] 188260 2.81%
#> (0.4,0.5] 88549 1.32%
#> (0.5,0.6] 42889 0.64%
#> (0.6,0.7] 21379 0.32%
#> (0.7,0.8] 11100 0.17%
#> (0.8,0.9] 6098 0.09%
#> (0.9,1] 2563 0.04%
#> [2026-07-26 22:27:35.698] [INFO ] [OK] Completed: create_edges [n_edges=6695970, n_comparisons=6695970, pass_rate=100.0%] (4.3s)
#> [2026-07-26 22:28:49.154] [INFO ] Selected community partition at resolution 0.025 (modularity 0.167)
#> [2026-07-26 22:28:51.064] [INFO ] Found 415 communities using weighted Louvain/Leiden clustering
#> [2026-07-26 22:28:51.320] [INFO ] > Starting: export_output [file=data/interim/features/example_edgesSpectra.tsv, n_rows=44188]
#> [2026-07-26 22:28:51.330] [INFO ] [OK] Completed: export_output [size_bytes=2234104] (10ms)
#> [2026-07-26 22:28:51.333] [INFO ] Edges written to: data/interim/features/example_edgesSpectra.tsv
#> [2026-07-26 22:28:51.334] [INFO ] [OK] Completed: create_edges_spectra [n_edges=44188, n_features=3660] (1m 34s)
#> ✔ fea_edg_spe completed [1m 34.1s, 2.23 MB]
#> + lib_rt_rts dispatched
#> ✔ lib_rt_rts completed [1ms, 86 B]
#> + lib_rt_sop dispatched
#> ✔ lib_rt_sop completed [1ms, 105 B]
#> + lib_spe_exp_int_pre_pos dispatched
#> ✔ lib_spe_exp_int_pre_pos completed [0ms, 600 B]
#> + lib_spe_exp_int_pre_neg dispatched
#> ✔ lib_spe_exp_int_pre_neg completed [0ms, 600 B]
#> + lib_spe_exp_int_pre_sop dispatched
#> ✔ lib_spe_exp_int_pre_sop completed [0ms, 79 B]
#> + fea_pre dispatched
#> [2026-07-26 22:28:54.296] [INFO ] > Starting: prepare_features_tables [input=data/source/example_features.csv, candidates=1]
#> [2026-07-26 22:28:54.635] [INFO ] Prepared 5328 feature-sample pairs
#> [2026-07-26 22:28:54.636] [INFO ] [OK] Completed: prepare_features_tables [n_features=5328] (340ms)
#> [2026-07-26 22:28:54.662] [INFO ] Exporting parameters to: data/interim/params/260726_222854_prepare_features_tables.yaml
#> [2026-07-26 22:28:54.664] [INFO ] > Starting: export_output [file=data/interim/features/example_features.tsv.gz, n_rows=5328]
#> [2026-07-26 22:28:54.691] [INFO ] [OK] Completed: export_output [size_bytes=172763] (27ms)
#> ✔ fea_pre completed [397ms, 172.76 kB]
#> + lib_sop_mer dispatched
#> [2026-07-26 22:28:55.146] [INFO ] > Starting: prepare_libraries_sop_merged [n_libraries=28, filter_enabled=FALSE, filter_level=none]
#> [2026-07-26 22:29:05.218] [INFO ] Splitting SOP library into standardized components
#> [2026-07-26 22:29:08.787] [INFO ] > Starting: process_smiles [n_structures=2203730]
#> [2026-07-26 22:29:08.789] [INFO ] Processing SMILES with RDKit
#> [2026-07-26 22:29:23.759] [INFO ] Processing 21 new SMILES with RDKit
#> [2026-07-26 22:29:23.761] [INFO ] Starting SMILES processing pipeline
#> [2026-07-26 22:29:23.761] [INFO ] Input: /tmp/RtmpZ8Qf0D/file26e51f534d58.smi
#> [2026-07-26 22:29:23.761] [INFO ] Output: /tmp/RtmpZ8Qf0D/file26e51298b2fc.csv.gz
#> [2026-07-26 22:29:23.761] [INFO ] Input file validated: /tmp/RtmpZ8Qf0D/file26e51f534d58.smi
#> [2026-07-26 22:29:23.761] [INFO ] Output file validated: /tmp/RtmpZ8Qf0D/file26e51298b2fc.csv.gz
#> [2026-07-26 22:29:23.761] [INFO ] Processing parameters: workers=8, batch_size=1000, progress_interval=10000
#> [2026-07-26 22:29:23.761] [INFO ] SMILES supplier initialized
#> [22:29:23] Explicit valence for atom # 1 N, 3, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 1
#> [22:29:23] ERROR: Explicit valence for atom # 1 N, 3, is greater than permitted
#> [22:29:23] Explicit valence for atom # 1 Cl, 7, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 4
#> [22:29:23] ERROR: Explicit valence for atom # 1 Cl, 7, is greater than permitted
#> [22:29:23] Explicit valence for atom # 1 Br, 3, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 5
#> [22:29:23] ERROR: Explicit valence for atom # 1 Br, 3, is greater than permitted
#> [22:29:23] Explicit valence for atom # 1 Br, 5, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 6
#> [22:29:23] ERROR: Explicit valence for atom # 1 Br, 5, is greater than permitted
#> [22:29:23] Explicit valence for atom # 1 Cl, 3, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 7
#> [22:29:23] ERROR: Explicit valence for atom # 1 Cl, 3, is greater than permitted
#> [22:29:23] Explicit valence for atom # 1 Cl, 5, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 8
#> [22:29:23] ERROR: Explicit valence for atom # 1 Cl, 5, is greater than permitted
#> [22:29:23] Explicit valence for atom # 1 I, 7, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 9
#> [22:29:23] ERROR: Explicit valence for atom # 1 I, 7, is greater than permitted
#> [22:29:23] Explicit valence for atom # 1 Cl, 3, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 10
#> [22:29:23] ERROR: Explicit valence for atom # 1 Cl, 3, is greater than permitted
#> [22:29:23] Explicit valence for atom # 8 Br, 2, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 13
#> [22:29:23] ERROR: Explicit valence for atom # 8 Br, 2, is greater than permitted
#> [22:29:23] Explicit valence for atom # 9 Cl, 2, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 14
#> [22:29:23] ERROR: Explicit valence for atom # 9 Cl, 2, is greater than permitted
#> [22:29:23] Explicit valence for atom # 6 C, 5, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 16
#> [22:29:23] ERROR: Explicit valence for atom # 6 C, 5, is greater than permitted
#> [22:29:23] Explicit valence for atom # 31 O, 3, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 17
#> [22:29:23] ERROR: Explicit valence for atom # 31 O, 3, is greater than permitted
#> [22:29:23] Explicit valence for atom # 4 N, 4, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 18
#> [22:29:23] ERROR: Explicit valence for atom # 4 N, 4, is greater than permitted
#> [22:29:23] Explicit valence for atom # 26 N, 4, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 19
#> [22:29:23] ERROR: Explicit valence for atom # 26 N, 4, is greater than permitted
#> [22:29:23] Explicit valence for atom # 0 P, 11, is greater than permitted
#> [22:29:23] ERROR: Could not sanitize molecule on line 20
#> [22:29:23] ERROR: Explicit valence for atom # 0 P, 11, is greater than permitted
#> [22:29:23] Can't kekulize mol. Unkekulized atoms: 6 7 8 9 10 11 12 13 14
#> [22:29:23] Can't kekulize mol. Unkekulized atoms: 6 7 8 9 10 11 12 13 14
#> [22:29:23] ERROR: Could not sanitize molecule on line 21
#> [22:29:23] ERROR: Can't kekulize mol. Unkekulized atoms: 6 7 8 9 10 11 12 13 14
#> [22:29:23] Explicit valence for atom # 56 P, 7, is greater than permitted
#> [2026-07-26 22:29:23.765] [WARNING] Failed to process SMILES 'CC(C)=CCCC(C)=CCCC(C)=CCCC(C)=CCCC(C)=CCCC(C)=CCCC(C)=CCCC(C)=CCCC(C)=CCCC(C)=CCCC(C)=CCO[P-]([O])(=O)=O': Explicit valence for atom # 56 P, 7, is greater than permitted
#> [22:29:23] Explicit valence for atom # 4 P, 7, is greater than permitted
#> [2026-07-26 22:29:23.765] [WARNING] Failed to process SMILES '[H][C@](O)(CO[P-]([O])(=O)=O)C=O': Explicit valence for atom # 4 P, 7, is greater than permitted
#> [22:29:23] Explicit valence for atom # 6 Si, 6, is greater than permitted
#> [2026-07-26 22:29:23.766] [WARNING] Failed to process SMILES 'C1=CC=C(C=C1)[Si-](C2=CC=CC=C2)(C3=CC=CC=C3)(F)F': Explicit valence for atom # 6 Si, 6, is greater than permitted
#> [22:29:23] Explicit valence for atom # 4 P, 7, is greater than permitted
#> [2026-07-26 22:29:23.767] [WARNING] Failed to process SMILES 'C(C(F)(F)[P-](C(C(F)(F)F)(F)F)(C(C(F)(F)F)(F)F)(F)(F)F)(F)(F)F': Explicit valence for atom # 4 P, 7, is greater than permitted
#> [22:29:23] Explicit valence for atom # 7 Si, 6, is greater than permitted
#> [2026-07-26 22:29:23.768] [WARNING] Failed to process SMILES 'C1=CC=C2C(=C1)O[Si-]3(O2)(OC4=CC=CC=C4O3)CI': Explicit valence for atom # 7 Si, 6, is greater than permitted
#> [2026-07-26 22:29:23.768] [WARNING] Batch processing: 5/5 molecules failed
#> [2026-07-26 22:29:23.768] [INFO ] Processing complete. Total molecules processed: 0
#> [2026-07-26 22:29:23.797] [INFO ] Successfully processed 0 SMILES
#> [2026-07-26 22:29:39.719] [INFO ] [OK] Completed: process_smiles [n_processed=2107823] (30.9s)
#> [2026-07-26 22:30:09.163] [INFO ] Referenced structure-organism pairs (1,328,980)
#> [2026-07-26 22:30:20.380] [INFO ] Structures: 430,553 stereoisomers, 1,376,522 without stereochemistry, 1,538,639 constitutional isomers
#> [2026-07-26 22:31:04.647] [INFO ] Unique organisms (37,469)
#> [2026-07-26 22:31:04.812] [INFO ] Processing 813 organism name(s) for OTT taxonomy lookup
#> [2026-07-26 22:31:05.204] [INFO ] Querying OTT API in 9 batches
#> [2026-07-26 22:31:09.642] [INFO ] Retrieving detailed taxonomy for 4 unique OTT IDs
#> [2026-07-26 22:31:10.379] [INFO ] Got OTTaxonomy!
#> [2026-07-26 22:31:10.910] [INFO ] Enriching NPClassifier taxonomy from additional cache: data/interim/libraries/sop/merged/structures/taxonomies/npc.tsv.gz
#> [2026-07-26 22:31:21.424] [INFO ] Enriched NPClassifier taxonomy with 1105925 entries from additional cache (1105925 missing keys matched)
#> [2026-07-26 22:31:27.845] [INFO ] Updated additional NPClassifier cache (1783925 total entries): data/interim/libraries/sop/merged/structures/taxonomies/npc.tsv.gz
#> [2026-07-26 22:31:28.009] [INFO ] Enriching ClassyFire taxonomy from additional cache: data/interim/libraries/sop/merged/structures/taxonomies/classyfire_cache.csv
#> [2026-07-26 22:31:33.565] [INFO ] Enriched ClassyFire taxonomy with 181659 entries from additional cache (181659 missing keys matched)
#> [2026-07-26 22:31:35.412] [INFO ] Updated additional ClassyFire cache (1106056 total entries): data/interim/libraries/sop/merged/structures/taxonomies/classyfire_cache.csv
#> [2026-07-26 22:31:35.439] [INFO ] Exporting parameters to: data/interim/params/260726_223135_prepare_libraries_sop_merged.yaml
#> [2026-07-26 22:31:35.441] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/merged/keys.tsv.gz, n_rows=1328980]
#> [2026-07-26 22:31:37.514] [INFO ] [OK] Completed: export_output [size_bytes=31597500] (2.1s)
#> [2026-07-26 22:31:37.516] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/merged/organisms/taxonomies/ott.tsv.gz, n_rows=36758]
#> [2026-07-26 22:31:37.610] [INFO ] [OK] Completed: export_output [size_bytes=1013371] (93ms)
#> [2026-07-26 22:31:37.612] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/merged/structures/canonical.tsv.gz, n_rows=2107823]
#> [2026-07-26 22:31:42.393] [INFO ] [OK] Completed: export_output [size_bytes=37376699] (4.8s)
#> [2026-07-26 22:31:42.395] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/merged/structures/stereo.tsv.gz, n_rows=1807075]
#> [2026-07-26 22:31:49.321] [INFO ] [OK] Completed: export_output [size_bytes=91160785] (6.9s)
#> [2026-07-26 22:31:49.323] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/merged/structures/metadata.tsv.gz, n_rows=1541988]
#> [2026-07-26 22:31:51.034] [INFO ] [OK] Completed: export_output [size_bytes=30595026] (1.7s)
#> [2026-07-26 22:31:51.035] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/merged/structures/taxonomies/classyfire.tsv.gz, n_rows=405842]
#> [2026-07-26 22:31:51.460] [INFO ] [OK] Completed: export_output [size_bytes=8389638] (425ms)
#> [2026-07-26 22:31:51.462] [INFO ] > Starting: export_output [file=data/interim/libraries/sop/merged/structures/taxonomies/npc.tsv.gz, n_rows=1326543]
#> [2026-07-26 22:31:53.693] [INFO ] [OK] Completed: export_output [size_bytes=19173400] (2.2s)
#> [2026-07-26 22:31:53.695] [INFO ] [OK] Completed: prepare_libraries_sop_merged [n_pairs=1328980, n_structures=1807075, n_organisms=36758, files_exported=7] (2m 59s)
#> ✔ lib_sop_mer completed [2m 58.6s, 219.31 MB]
#> + lib_mer_str_met dispatched
#> ✔ lib_mer_str_met completed [0ms, 30.60 MB]
#> + lib_mer_org_tax_ott dispatched
#> ✔ lib_mer_org_tax_ott completed [1ms, 1.01 MB]
#> + lib_mer_key dispatched
#> ✔ lib_mer_key completed [0ms, 31.60 MB]
#> + lib_mer_str_tax_npc dispatched
#> ✔ lib_mer_str_tax_npc completed [0ms, 19.17 MB]
#> + lib_mer_str_stereo dispatched
#> ✔ lib_mer_str_stereo completed [1ms, 91.16 MB]
#> + lib_mer_str_tax_cla dispatched
#> ✔ lib_mer_str_tax_cla completed [0ms, 8.39 MB]
#> + tax_pre dispatched
#> [2026-07-26 22:31:58.684] [INFO ] > Starting: prepare_taxa [taxon=NULL]
#> [2026-07-26 22:31:58.855] [INFO ] Processing 2 organism name(s) for OTT taxonomy lookup
#> [2026-07-26 22:31:59.120] [INFO ] Querying OTT API in 1 batches
#> [2026-07-26 22:31:59.313] [INFO ] Retrying failed queries using genus names only
#> [2026-07-26 22:31:59.320] [INFO ] Retrying with 1 genus names: blk
#> [2026-07-26 22:31:59.513] [INFO ] Retrieving detailed taxonomy for 1 unique OTT IDs
#> [2026-07-26 22:31:59.632] [INFO ] Got OTTaxonomy!
#> [2026-07-26 22:32:00.086] [INFO ] [OK] Completed: prepare_taxa [n_features=5328] (1.4s)
#> [2026-07-26 22:32:00.119] [INFO ] Exporting parameters to: data/interim/params/260726_223200_prepare_taxa.yaml
#> [2026-07-26 22:32:00.121] [INFO ] > Starting: export_output [file=data/interim/taxa/example_taxed.tsv.gz, n_rows=5328]
#> [2026-07-26 22:32:00.128] [INFO ] [OK] Completed: export_output [size_bytes=19697] (7ms)
#> ✔ tax_pre completed [1.4s, 19.70 kB]
#> + ann_sir_pre dispatched
#> [2026-07-26 22:32:00.607] [INFO ] > Starting: prepare_annotations_sirius [version=6]
#> [2026-07-26 22:32:00.773] [INFO ] > Starting: process_smiles [n_structures=2563]
#> [2026-07-26 22:32:00.774] [INFO ] Processing SMILES with RDKit
#> [2026-07-26 22:32:08.944] [INFO ] Processing 9 new SMILES with RDKit
#> [2026-07-26 22:32:08.946] [INFO ] Starting SMILES processing pipeline
#> [2026-07-26 22:32:08.946] [INFO ] Input: /tmp/RtmpZ8Qf0D/file26e53a210a45.smi
#> [2026-07-26 22:32:08.946] [INFO ] Output: /tmp/RtmpZ8Qf0D/file26e53bebf02b.csv.gz
#> [2026-07-26 22:32:08.946] [INFO ] Input file validated: /tmp/RtmpZ8Qf0D/file26e53a210a45.smi
#> [2026-07-26 22:32:08.946] [INFO ] Output file validated: /tmp/RtmpZ8Qf0D/file26e53bebf02b.csv.gz
#> [2026-07-26 22:32:08.946] [INFO ] Processing parameters: workers=8, batch_size=1000, progress_interval=10000
#> [2026-07-26 22:32:08.946] [INFO ] SMILES supplier initialized
#> [22:32:08] Explicit valence for atom # 8 Cl, 3, is greater than permitted
#> [22:32:08] ERROR: Could not sanitize molecule on line 1
#> [22:32:08] ERROR: Explicit valence for atom # 8 Cl, 3, is greater than permitted
#> [22:32:08] Explicit valence for atom # 4 P, 7, is greater than permitted
#> [22:32:08] ERROR: Could not sanitize molecule on line 2
#> [22:32:08] ERROR: Explicit valence for atom # 4 P, 7, is greater than permitted
#> [22:32:08] Explicit valence for atom # 2 P, 7, is greater than permitted
#> [22:32:08] ERROR: Could not sanitize molecule on line 3
#> [22:32:08] ERROR: Explicit valence for atom # 2 P, 7, is greater than permitted
#> [22:32:08] Explicit valence for atom # 4 P, 7, is greater than permitted
#> [22:32:08] ERROR: Could not sanitize molecule on line 4
#> [22:32:08] ERROR: Explicit valence for atom # 4 P, 7, is greater than permitted
#> [22:32:08] Explicit valence for atom # 2 P, 7, is greater than permitted
#> [22:32:08] ERROR: Could not sanitize molecule on line 5
#> [22:32:08] ERROR: Explicit valence for atom # 2 P, 7, is greater than permitted
#> [22:32:08] Explicit valence for atom # 6 P, 7, is greater than permitted
#> [22:32:08] ERROR: Could not sanitize molecule on line 6
#> [22:32:08] ERROR: Explicit valence for atom # 6 P, 7, is greater than permitted
#> [22:32:08] Explicit valence for atom # 6 P, 7, is greater than permitted
#> [22:32:08] ERROR: Could not sanitize molecule on line 7
#> [22:32:08] ERROR: Explicit valence for atom # 6 P, 7, is greater than permitted
#> [22:32:08] Explicit valence for atom # 4 P, 7, is greater than permitted
#> [22:32:08] ERROR: Could not sanitize molecule on line 8
#> [22:32:08] ERROR: Explicit valence for atom # 4 P, 7, is greater than permitted
#> [22:32:08] Explicit valence for atom # 2 P, 7, is greater than permitted
#> [22:32:08] ERROR: Could not sanitize molecule on line 9
#> [22:32:08] ERROR: Explicit valence for atom # 2 P, 7, is greater than permitted
#> [2026-07-26 22:32:08.947] [INFO ] Processing complete. Total molecules processed: 0
#> [2026-07-26 22:32:08.977] [INFO ] Successfully processed 0 SMILES
#> [2026-07-26 22:32:17.349] [INFO ] [OK] Completed: process_smiles [n_processed=2554] (16.6s)
#> [2026-07-26 22:32:17.371] [INFO ] > Starting: complement_metadata [n_input=2571]
#> [2026-07-26 22:33:07.912] [INFO ] [OK] Completed: complement_metadata [n_enriched=2571] (50.5s)
#> [2026-07-26 22:33:07.924] [INFO ] [OK] Completed: prepare_annotations_sirius [n_canopus=15, n_formulas=19, n_structures=2571] (1m 7s)
#> [2026-07-26 22:33:07.952] [INFO ] Exporting parameters to: data/interim/params/260726_223307_prepare_annotations_sirius.yaml
#> [2026-07-26 22:33:07.953] [INFO ] > Starting: export_output [file=data/interim/annotations/example_canopusPrepared.tsv.gz, n_rows=15]
#> [2026-07-26 22:33:07.955] [INFO ] [OK] Completed: export_output [size_bytes=830] (1ms)
#> [2026-07-26 22:33:07.956] [INFO ] > Starting: export_output [file=data/interim/annotations/example_formulaPrepared.tsv.gz, n_rows=19]
#> [2026-07-26 22:33:07.957] [INFO ] [OK] Completed: export_output [size_bytes=521] (1ms)
#> [2026-07-26 22:33:07.959] [INFO ] > Starting: export_output [file=data/interim/annotations/example_siriusPrepared.tsv.gz, n_rows=2571]
#> [2026-07-26 22:33:07.973] [INFO ] [OK] Completed: export_output [size_bytes=98016] (14ms)
#> ✔ ann_sir_pre completed [1m 7.4s, 99.37 kB]
#> + ann_ms1_pre dispatched
#> [2026-07-26 22:33:09.807] [INFO ] > Starting: annotate_masses [ms_mode=pos, tolerance_ppm=10, tolerance_dalton=0.005, tolerance_rt=0.02]
#> [2026-07-26 22:33:09.809] [INFO ] Starting mass-based annotation (streaming network-first path)
#> [2026-07-26 22:33:09.810] [INFO ] ============================================================
#> [2026-07-26 22:33:09.811] [INFO ] Data Sanitizing: Pre-flight Checks
#> [2026-07-26 22:33:09.812] [INFO ] ============================================================
#> [2026-07-26 22:33:09.813] [INFO ] Checking features file...
#> [2026-07-26 22:33:09.858] [INFO ] [OK] Features file: 5328 rows, 7 columns
#> [2026-07-26 22:33:09.859] [INFO ] ============================================================
#> [2026-07-26 22:33:09.860] [INFO ] [OK] All pre-flight checks passed!
#> [2026-07-26 22:33:09.861] [INFO ] Data validation complete. Ready to proceed.
#> [2026-07-26 22:33:09.862] [INFO ] ============================================================
#> [2026-07-26 22:33:09.901] [INFO ] Processing 5328 features for annotation
#> [2026-07-26 22:33:09.902] [INFO ] > Starting: harmonize_adducts [n_rows=5328]
#> [2026-07-26 22:33:09.928] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=13, n_unique_after=13] (25ms)
#> [2026-07-26 22:33:09.946] [INFO ] > Starting: harmonize_adducts [n_rows=2112]
#> [2026-07-26 22:33:09.947] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=13, n_unique_after=13] (1ms)
#> [2026-07-26 22:33:09.948] [INFO ] Pre-assigned adducts kept as hypotheses alongside the [M+H]+ baseline: 2112
#> [2026-07-26 22:33:36.301] [INFO ] Built 22194 RT-window pair(s)
#> [2026-07-26 22:33:36.557] [INFO ] Here are the top 16 observed m/z differences inside the RT windows:
#> [2026-07-26 22:33:36.559] [INFO ]
#> bin N Pct
#> (4.94331,4.96196] 316 18.93%
#> (21.9716,21.9902] 247 14.80%
#> (17.0104,17.0291] 174 10.43%
#> (17.9989,18.0176] 142 8.51%
#> (38.9998,39.0185] 114 6.83%
#> (15.966,15.9846] 105 6.29%
#> (39.9883,40.007] 96 5.75%
#> (77.9989,78.0175] 84 5.03%
#> (18.4839,18.5025] 74 4.43%
#> (35.0272,35.0459] 72 4.31%
#> (0.989321,1.00797] 61 3.65%
#> (162.04,162.058] 38 2.28%
#> (30.0101,30.0288] 38 2.28%
#> (28.0145,28.0331] 37 2.22%
#> (2.01512,2.03377] 37 2.22%
#> (109.948,109.967] 34 2.04%
#> [2026-07-26 22:33:37.930] [INFO ] Evidence engine: 5328 features x 24 adducts (prefilter=on, cap=Inf)
#> [2026-07-26 22:33:37.987] [INFO ] Evidence engine candidate materialization: 92276 rows
#> [2026-07-26 22:33:43.295] [INFO ] Evidence engine complete: 7492 rows, 5710 supported clusters
#> [2026-07-26 22:33:55.248] [INFO ] > Starting: harmonize_adducts [n_rows=5328]
#> [2026-07-26 22:33:55.250] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=6, n_unique_after=6] (2ms)
#> [2026-07-26 22:33:59.168] [INFO ] Pairwise-support filter removed 81 modifier-bearing evidence hypothesis row(s) lacking direct adduct/cluster/loss support.
#> [2026-07-26 22:33:59.191] [INFO ] Evidence-based discovery added 1324 adduct edge(s)
#> [2026-07-26 22:34:39.201] [INFO ] Intensity co-variance: tested 85477 pair(s), retained 736 significant edge(s) (p < 0.05); r in [0.88, 1.00], median r = 1.00. Dropped: 0 not in matrix, 80508 insufficient samples, 0 correlation failed, 2275 non-positive, 1958 p >= 0.05
#> [2026-07-26 22:34:39.238] [INFO ] Edge classification complete in 62.56 seconds: 1722 adduct edges, 736 covariance edges, 53 cluster edges, 830 loss edges
#> [2026-07-26 22:34:46.690] [INFO ] > Starting: harmonize_adducts [n_rows=23676]
#> [2026-07-26 22:34:46.693] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=25, n_unique_after=25] (4ms)
#> [2026-07-26 22:34:48.160] [INFO ] > Starting: harmonize_adducts [n_rows=1267]
#> [2026-07-26 22:34:48.515] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=268, n_unique_after=262] (354ms)
#> [2026-07-26 22:35:09.085] [INFO ] Constrained multi-adduct expansion kept 1072 hypothesis row(s)
#> [2026-07-26 22:35:10.365] [INFO ] Network-consensus pruning dropped 105 (feature, adduct) candidate(s) with zero adduct-graph support when a supported alternative existed.
#> [2026-07-26 22:35:34.228] [INFO ] Loss/formula filter demoted 12856 structural match row(s) where the actual loss-edge term was not contained in the candidate formula.
#> [2026-07-26 22:35:34.735] [INFO ] Annotation/edge adduct agreement removed 20 unsupported (feature, adduct) assignment(s).
#> [2026-07-26 22:35:55.433] [INFO ] Conflict-resolution filter removed 5696 annotation row(s) with states incompatible with graph-consistent evidence.
#> [2026-07-26 22:35:55.435] [INFO ] Conflict-resolution pruning touched 3184 feature(s) and removed all annotations from 47 feature(s).
#> [2026-07-26 22:35:55.590] [INFO ] Library matching complete in 45.80 seconds: 610773 annotations
#> [2026-07-26 22:36:32.104] [INFO ] Coverage audit: kept 610773/670469 annotation rows across 5281/5328 features; 59696 annotation rows were pruned from 3184 feature(s).
#> [2026-07-26 22:36:32.106] [INFO ] > Starting: decorate_masses [n_annotations=610773]
#> [2026-07-26 22:36:32.294] [INFO ] MS1 annotations: 238973 unique structures across 5195 features
#> [2026-07-26 22:36:32.295] [INFO ] [OK] Completed: decorate_masses [n_structures=238973, n_features=5195] (189ms)
#> [2026-07-26 22:36:32.606] [INFO ] Breakdown of the (top 64) annotated adduct species (library-matched):
#> [2026-07-26 22:36:32.610] [INFO ]
#> adduct N_features N_annotations Pct_features Pct_annotations
#> [M+Na]+ 3285 125382 19.88% 20.53%
#> [M+H]+ 3134 165148 18.96% 27.04%
#> [M+H4N]+ 3130 103222 18.94% 16.90%
#> [M+K]+ 2579 75830 15.61% 12.42%
#> [M+H2]2+ 2008 26533 12.15% 4.35%
#> [2M+H]+ 1520 72895 9.20% 11.94%
#> [M-H2O+H]+ 148 6783 0.90% 1.11%
#> [M+H3N+H]+ 48 543 0.29% 0.09%
#> [2M+H4N]+ 46 3425 0.28% 0.56%
#> [2M+Na]+ 43 4789 0.26% 0.78%
#> [2M-H2O+H]+ 37 1352 0.22% 0.22%
#> [M-H+2Na]+ 36 1577 0.22% 0.26%
#> [M+Ca]2+ 36 549 0.22% 0.09%
#> [M]+ 26 614 0.16% 0.10%
#> [M+C2H4+H4N]+ 24 2183 0.15% 0.36%
#> [2M+Ca]2+ 17 904 0.10% 0.15%
#> [M-C6H12O6+H]+ 14 250 0.08% 0.04%
#> [2M+C2H4+H4N]+ 10 742 0.06% 0.12%
#> [2M+K]+ 9 339 0.05% 0.06%
#> [M-H4O2+H]+ 9 293 0.05% 0.05%
#> [M-C6H10O5+H]+ 9 258 0.05% 0.04%
#> [M+2H]2+ 8 219 0.05% 0.04%
#> [M+C3H4O+Na]+ 8 171 0.05% 0.03%
#> [M+Cu]+ 7 150 0.04% 0.02%
#> [2M+Fe]2+ 7 138 0.04% 0.02%
#> [M+C2H3N+Na]+ 5 1109 0.03% 0.18%
#> [2M+Mg]2+ 5 306 0.03% 0.05%
#> [M-C6H14O7+H]+ 5 124 0.03% 0.02%
#> [M+O+Na]+ 4 672 0.02% 0.11%
#> [M+Fe]2+ 4 58 0.02% 0.01%
#> [M+CHO2+K]+ 4 50 0.02% 0.01%
#> [M+C6H10O4+H]+ 3 388 0.02% 0.06%
#> [M+CH2O+H]+ 3 331 0.02% 0.05%
#> [M+CH2O+Na]+ 3 320 0.02% 0.05%
#> [M+H2O2+Na]+ 3 237 0.02% 0.04%
#> [M-C6H10O4+H]+ 3 174 0.02% 0.03%
#> [M-H2O2+H]+ 3 158 0.02% 0.03%
#> [M+C4H4O+H4N]+ 3 132 0.02% 0.02%
#> [2M+CHO2+K]+ 3 125 0.02% 0.02%
#> [M-H6O3+H]+ 3 123 0.02% 0.02%
#> [2M-H4O2+H]+ 3 102 0.02% 0.02%
#> [M-C4H4O+Na]+ 3 88 0.02% 0.01%
#> [M-C6H10O4+H4N]+ 3 48 0.02% 0.01%
#> [M-C5H8O4+H]+ 3 40 0.02% 0.01%
#> [2M+O+Na]+ 2 360 0.01% 0.06%
#> [2M-CO+H]+ 2 284 0.01% 0.05%
#> [M-CO+H]+ 2 203 0.01% 0.03%
#> [M-C3O3+Na]+ 2 188 0.01% 0.03%
#> [M-CH2O+H]+ 2 170 0.01% 0.03%
#> [2M+C3H6O2+H]+ 2 161 0.01% 0.03%
#> [2M-C2H4+H4N]+ 2 145 0.01% 0.02%
#> [M-C2H4+H]+ 2 145 0.01% 0.02%
#> [M-CH2O+Na]+ 2 143 0.01% 0.02%
#> [M-CHNO+H]+ 2 133 0.01% 0.02%
#> [M+C2H2O3+Na]+ 2 120 0.01% 0.02%
#> [M-C2H2O+Na]+ 2 117 0.01% 0.02%
#> [M-O+H4N]+ 2 91 0.01% 0.01%
#> [M-C6H12O5+H]+ 2 84 0.01% 0.01%
#> [M-C2H4O2+H]+ 2 80 0.01% 0.01%
#> [M-C7H4O4+Na]+ 2 75 0.01% 0.01%
#> [M-CHNO+H4N]+ 2 61 0.01% 0.01%
#> [M-C6H6O3+Na]+ 2 60 0.01% 0.01%
#> [M+C3H4O3+H4N]+ 2 49 0.01% 0.01%
#> [M-H2O+Na]+ 2 49 0.01% 0.01%
#> [2026-07-26 22:36:32.618] [INFO ] Adduct hypotheses retained without library match (by source):
#> [2026-07-26 22:36:32.619] [INFO ]
#> source N_features N_adduct_types
#> loss 44 2
#> pair 34 6
#> preassigned 21 7
#> evidence 18 8
#> baseline 13 1
#> [2026-07-26 22:36:35.152] [INFO ] Exporting parameters to: data/interim/params/260726_223635_annotate_masses.yaml
#> [2026-07-26 22:36:35.154] [INFO ] > Starting: export_output [file=data/interim/features/example_edgesMasses.tsv, n_rows=5150]
#> [2026-07-26 22:36:35.156] [INFO ] [OK] Completed: export_output [size_bytes=129515] (2ms)
#> [2026-07-26 22:36:35.157] [INFO ] Exported edges: example_edgesMasses.tsv (5,150 rows)
#> [2026-07-26 22:36:39.268] [INFO ] > Starting: process_smiles [n_structures=239014]
#> [2026-07-26 22:36:39.270] [INFO ] Processing SMILES with RDKit
#> [2026-07-26 22:36:50.814] [INFO ] Processing 76 new SMILES with RDKit
#> [2026-07-26 22:36:50.816] [INFO ] Starting SMILES processing pipeline
#> [2026-07-26 22:36:50.816] [INFO ] Input: /tmp/RtmpZ8Qf0D/file26e52d5fb5ae.smi
#> [2026-07-26 22:36:50.816] [INFO ] Output: /tmp/RtmpZ8Qf0D/file26e5554f65f0.csv.gz
#> [2026-07-26 22:36:50.816] [INFO ] Input file validated: /tmp/RtmpZ8Qf0D/file26e52d5fb5ae.smi
#> [2026-07-26 22:36:50.816] [INFO ] Output file validated: /tmp/RtmpZ8Qf0D/file26e5554f65f0.csv.gz
#> [2026-07-26 22:36:50.816] [INFO ] Processing parameters: workers=8, batch_size=1000, progress_interval=10000
#> [2026-07-26 22:36:50.816] [INFO ] SMILES supplier initialized
#> [2026-07-26 22:36:50.950] [INFO ] Processing complete. Total molecules processed: 76
#> [2026-07-26 22:36:50.984] [INFO ] Successfully processed 76 SMILES
#> [2026-07-26 22:36:59.868] [INFO ] [OK] Completed: process_smiles [n_processed=239014] (20.6s)
#> [2026-07-26 22:37:02.194] [INFO ] > Starting: complement_metadata [n_input=610773]
#> [2026-07-26 22:37:20.632] [INFO ] [OK] Completed: complement_metadata [n_enriched=610773] (18.4s)
#> [2026-07-26 22:37:20.636] [INFO ] > Starting: export_output [file=data/interim/annotations/example_ms1Prepared.tsv.gz, n_rows=610773]
#> [2026-07-26 22:37:23.469] [INFO ] [OK] Completed: export_output [size_bytes=42942124] (2.8s)
#> [2026-07-26 22:37:23.470] [INFO ] Exported annotations: example_ms1Prepared.tsv.gz (610,773 rows)
#> [2026-07-26 22:37:23.472] [INFO ] > Starting: export_output [file=data/interim/annotations/example_ms1Prepared_coverage.tsv.gz, n_rows=11]
#> [2026-07-26 22:37:23.473] [INFO ] [OK] Completed: export_output [size_bytes=283] (1ms)
#> [2026-07-26 22:37:23.475] [INFO ] Exported coverage report: example_ms1Prepared_coverage.tsv.gz
#> [2026-07-26 22:37:23.476] [INFO ] All outputs exported in 50.86 seconds
#> [2026-07-26 22:37:23.477] [INFO ] [OK] Completed: annotate_masses [n_annotations=610773, n_edges=5150] (4m 14s)
#> ✔ ann_ms1_pre completed [4m 13.7s, 43.07 MB]
#> + ann_spe_exp_gnp_pre dispatched
#> [2026-07-26 22:37:25.926] [INFO ] > Starting: prepare_annotations_gnps [n_files=1]
#> [2026-07-26 22:37:25.929] [WARN ] No GNPS annotations found, returning an empty file instead
#> [2026-07-26 22:37:25.931] [INFO ] [OK] Completed: prepare_annotations_gnps [n_annotations=1] (6ms)
#> [2026-07-26 22:37:25.951] [INFO ] Exporting parameters to: data/interim/params/260726_223725_prepare_annotations_gnps.yaml
#> [2026-07-26 22:37:25.953] [INFO ] > Starting: export_output [file=data/interim/annotations/example_gnpsPrepared.tsv.gz, n_rows=1]
#> [2026-07-26 22:37:25.954] [INFO ] [OK] Completed: export_output [size_bytes=308] (1ms)
#> ✔ ann_spe_exp_gnp_pre completed [35ms, 308 B]
#> + ann_spe_exp_mzm_pre dispatched
#> [2026-07-26 22:37:27.740] [INFO ] > Starting: prepare_annotations_mzmine [n_files=1]
#> [2026-07-26 22:37:27.742] [WARN ] No mzmine annotations found, returning an empty file instead
#> [2026-07-26 22:37:27.744] [INFO ] [OK] Completed: prepare_annotations_mzmine [n_annotations=1] (4ms)
#> [2026-07-26 22:37:27.760] [INFO ] Exporting parameters to: data/interim/params/260726_223727_prepare_annotations_mzmine.yaml
#> [2026-07-26 22:37:27.761] [INFO ] > Starting: export_output [file=data/interim/annotations/example_mzminePrepared.tsv.gz, n_rows=1]
#> [2026-07-26 22:37:27.762] [INFO ] [OK] Completed: export_output [size_bytes=308] (1ms)
#> ✔ ann_spe_exp_mzm_pre completed [25ms, 308 B]
#> + ann_spe_exp_mzt_pre dispatched
#> [2026-07-26 22:37:29.531] [WARN ] No mzTab input provided for prepare_annotations_mztab, exporting empty annotations
#> [2026-07-26 22:37:29.535] [INFO ] > Starting: export_output [file=data/interim/annotations/example_mztabPrepared.tsv.gz, n_rows=1]
#> [2026-07-26 22:37:29.536] [INFO ] [OK] Completed: export_output [size_bytes=308] (1ms)
#> ✔ ann_spe_exp_mzt_pre completed [8ms, 308 B]
#> + ann_sir_pre_can dispatched
#> ✔ ann_sir_pre_can completed [0ms, 830 B]
#> + ann_sir_pre_for dispatched
#> ✔ ann_sir_pre_for completed [1ms, 521 B]
#> + ann_sir_pre_str dispatched
#> ✔ ann_sir_pre_str completed [0ms, 98.02 kB]
#> + ann_ms1_pre_ann dispatched
#> ✔ ann_ms1_pre_ann completed [0ms, 42.94 MB]
#> + ann_ms1_pre_edg dispatched
#> ✔ ann_ms1_pre_edg completed [0ms, 129.51 kB]
#> + ann_spe_pos dispatched
#> [2026-07-26 22:37:40.296] [INFO ] ============================================================
#> [2026-07-26 22:37:40.297] [INFO ] Data Sanitizing: Pre-flight Checks
#> [2026-07-26 22:37:40.298] [INFO ] ============================================================
#> [2026-07-26 22:37:40.299] [INFO ] Checking MGF file...
#> [2026-07-26 22:37:40.820] [INFO ] [OK] MGF file: 12195 MS2 spectra found
#> [2026-07-26 22:37:40.821] [INFO ] ============================================================
#> [2026-07-26 22:37:40.822] [INFO ] [OK] All pre-flight checks passed!
#> [2026-07-26 22:37:40.823] [INFO ] Data validation complete. Ready to proceed.
#> [2026-07-26 22:37:40.824] [INFO ] ============================================================
#> [2026-07-26 22:37:40.825] [INFO ] Starting spectral annotation in pos mode
#> [2026-07-26 22:37:40.827] [INFO ] Importing spectra from: data/source/example_spectra.mgf
#> [2026-07-26 22:37:40.828] [INFO ] Reading MGF file (7.41 MB) with optimized parser: data/source/example_spectra.mgf
#> [2026-07-26 22:37:43.179] [INFO ] Processed 10000 spectra...
#> [2026-07-26 22:37:44.876] [INFO ] Total spectra read: 16282
#> [2026-07-26 22:37:52.784] [INFO ] Loaded 16282 spectra from file
#> [2026-07-26 22:37:52.803] [INFO ] Combining replicate spectra by FEATURE_ID
#> [2026-07-26 22:37:54.041] [INFO ] Combined replicates: 12195 -> 4087 spectra
#> [2026-07-26 22:37:54.078] [INFO ] Sanitizing 4087 spectra (cutoff: 0)
#> [2026-07-26 22:37:55.634] [INFO ] Sanitization complete: 3660/4087 spectra retained (89.6%, 427 removed)
#> [2026-07-26 22:37:55.636] [INFO ] Import complete: 3660 spectra ready for analysis
#> [2026-07-26 22:37:55.638] [INFO ] > Starting: harmonize_adducts [n_rows=3660]
#> [2026-07-26 22:37:55.639] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=12, n_unique_after=12] (2ms)
#> [2026-07-26 22:37:58.298] [INFO ] > Starting: harmonize_adducts [n_rows=610773]
#> [2026-07-26 22:37:58.356] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=272, n_unique_after=272] (58ms)
#> [2026-07-26 22:38:00.509] [INFO ] > Starting: harmonize_adducts [n_rows=210419]
#> [2026-07-26 22:38:00.525] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=2, n_unique_after=2] (16ms)
#> [2026-07-26 22:38:00.891] [INFO ]
#> library spectra unique_structures Pct_spectra
#> ISDB - NormanSusDat 81390 33748 100.00%
#> [2026-07-26 22:38:00.894] [INFO ] > Starting: calculate_entropy_similarity [n_library=81390, n_query=3660, method=gnps]
#> [2026-07-26 22:38:00.895] [INFO ] Calculating entropy and similarity for 3660 spectra
#> [2026-07-26 22:38:02.860] [INFO ] Processed 500 / 3660 queries
#> [2026-07-26 22:38:04.405] [INFO ] Processed 1000 / 3660 queries
#> [2026-07-26 22:38:05.302] [INFO ] Processed 1500 / 3660 queries
#> [2026-07-26 22:38:06.671] [INFO ] Processed 2000 / 3660 queries
#> [2026-07-26 22:38:07.356] [INFO ] Processed 2500 / 3660 queries
#> [2026-07-26 22:38:08.471] [INFO ] Processed 3000 / 3660 queries
#> [2026-07-26 22:38:09.033] [INFO ] Processed 3500 / 3660 queries
#> [2026-07-26 22:38:09.174] [INFO ] Processed 3660 / 3660 queries
#> [2026-07-26 22:38:09.196] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=93005] (8.3s)
#> [2026-07-26 22:38:09.200] [INFO ] > Starting: harmonize_adducts [n_rows=81390]
#> [2026-07-26 22:38:09.207] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=2, n_unique_after=2] (7ms)
#> [2026-07-26 22:38:09.324] [INFO ] > Starting: calculate_entropy_similarity [n_library=59920, n_query=3644, method=gnps]
#> [2026-07-26 22:38:09.329] [INFO ] Calculating entropy and similarity for 3644 spectra
#> [2026-07-26 22:38:11.949] [INFO ] Processed 500 / 3644 queries
#> [2026-07-26 22:38:13.079] [INFO ] Processed 1000 / 3644 queries
#> [2026-07-26 22:38:14.605] [INFO ] Processed 1500 / 3644 queries
#> [2026-07-26 22:38:16.057] [INFO ] Processed 2000 / 3644 queries
#> [2026-07-26 22:38:17.417] [INFO ] Processed 2500 / 3644 queries
#> [2026-07-26 22:38:18.143] [INFO ] Processed 3000 / 3644 queries
#> [2026-07-26 22:38:19.312] [INFO ] Processed 3500 / 3644 queries
#> [2026-07-26 22:38:19.493] [INFO ] Processed 3644 / 3644 queries
#> [2026-07-26 22:38:19.518] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=109826] (10.2s)
#> [2026-07-26 22:38:19.548] [INFO ] > Starting: harmonize_adducts [n_rows=81360]
#> [2026-07-26 22:38:19.556] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=2, n_unique_after=2] (8ms)
#> [2026-07-26 22:38:37.956] [INFO ] > Starting: harmonize_adducts [n_rows=994408]
#> [2026-07-26 22:38:38.019] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=1, n_unique_after=1] (64ms)
#> [2026-07-26 22:38:42.265] [INFO ]
#> library spectra unique_structures Pct_spectra
#> ISDB - Wikidata 442120 442119 100.00%
#> [2026-07-26 22:38:42.266] [INFO ] > Starting: calculate_entropy_similarity [n_library=442120, n_query=3660, method=gnps]
#> [2026-07-26 22:38:42.267] [INFO ] Calculating entropy and similarity for 3660 spectra
#> [2026-07-26 22:38:56.565] [INFO ] Processed 500 / 3660 queries
#> [2026-07-26 22:39:06.729] [INFO ] Processed 1000 / 3660 queries
#> [2026-07-26 22:39:15.252] [INFO ] Processed 1500 / 3660 queries
#> [2026-07-26 22:39:22.812] [INFO ] Processed 2000 / 3660 queries
#> [2026-07-26 22:39:29.596] [INFO ] Processed 2500 / 3660 queries
#> [2026-07-26 22:39:35.859] [INFO ] Processed 3000 / 3660 queries
#> [2026-07-26 22:39:40.891] [INFO ] Processed 3500 / 3660 queries
#> [2026-07-26 22:39:42.364] [INFO ] Processed 3660 / 3660 queries
#> [2026-07-26 22:39:42.400] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=553373] (1m)
#> [2026-07-26 22:39:42.406] [INFO ] > Starting: harmonize_adducts [n_rows=442120]
#> [2026-07-26 22:39:42.436] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=1, n_unique_after=1] (30ms)
#> [2026-07-26 22:39:42.691] [INFO ] > Starting: calculate_entropy_similarity [n_library=322054, n_query=3644, method=gnps]
#> [2026-07-26 22:39:42.710] [INFO ] Calculating entropy and similarity for 3644 spectra
#> [2026-07-26 22:39:59.518] [INFO ] Processed 500 / 3644 queries
#> [2026-07-26 22:40:12.515] [INFO ] Processed 1000 / 3644 queries
#> [2026-07-26 22:40:23.622] [INFO ] Processed 1500 / 3644 queries
#> [2026-07-26 22:40:30.926] [INFO ] Processed 2000 / 3644 queries
#> [2026-07-26 22:40:37.825] [INFO ] Processed 2500 / 3644 queries
#> [2026-07-26 22:40:43.434] [INFO ] Processed 3000 / 3644 queries
#> [2026-07-26 22:40:48.327] [INFO ] Processed 3500 / 3644 queries
#> [2026-07-26 22:40:49.236] [INFO ] Processed 3644 / 3644 queries
#> [2026-07-26 22:40:49.274] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=598956] (1m 7s)
#> [2026-07-26 22:40:49.420] [INFO ] > Starting: harmonize_adducts [n_rows=442025]
#> [2026-07-26 22:40:49.449] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=1, n_unique_after=1] (29ms)
#> [2026-07-26 22:41:13.227] [INFO ] > Starting: harmonize_adducts [n_rows=891903]
#> [2026-07-26 22:41:13.287] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=6, n_unique_after=6] (60ms)
#> [2026-07-26 22:41:15.434] [INFO ]
#> library spectra unique_structures Pct_spectra
#> enveda180 549399 118497 100.00%
#> [2026-07-26 22:41:15.435] [INFO ] > Starting: calculate_entropy_similarity [n_library=549399, n_query=3660, method=gnps]
#> [2026-07-26 22:41:15.437] [INFO ] Calculating entropy and similarity for 3660 spectra
#> [2026-07-26 22:41:30.347] [INFO ] Processed 500 / 3660 queries
#> [2026-07-26 22:41:40.422] [INFO ] Processed 1000 / 3660 queries
#> [2026-07-26 22:41:49.305] [INFO ] Processed 1500 / 3660 queries
#> [2026-07-26 22:41:57.415] [INFO ] Processed 2000 / 3660 queries
#> [2026-07-26 22:42:04.920] [INFO ] Processed 2500 / 3660 queries
#> [2026-07-26 22:42:11.338] [INFO ] Processed 3000 / 3660 queries
#> [2026-07-26 22:42:17.031] [INFO ] Processed 3500 / 3660 queries
#> [2026-07-26 22:42:18.906] [INFO ] Processed 3660 / 3660 queries
#> [2026-07-26 22:42:18.933] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=644031] (1m 3s)
#> [2026-07-26 22:42:18.939] [INFO ] > Starting: harmonize_adducts [n_rows=549399]
#> [2026-07-26 22:42:18.977] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=6, n_unique_after=6] (37ms)
#> [2026-07-26 22:42:19.307] [INFO ] > Starting: calculate_entropy_similarity [n_library=418241, n_query=3644, method=gnps]
#> [2026-07-26 22:42:19.332] [INFO ] Calculating entropy and similarity for 3644 spectra
#> [2026-07-26 22:42:36.185] [INFO ] Processed 500 / 3644 queries
#> [2026-07-26 22:42:48.227] [INFO ] Processed 1000 / 3644 queries
#> [2026-07-26 22:42:59.437] [INFO ] Processed 1500 / 3644 queries
#> [2026-07-26 22:43:10.786] [INFO ] Processed 2000 / 3644 queries
#> [2026-07-26 22:43:20.551] [INFO ] Processed 2500 / 3644 queries
#> [2026-07-26 22:43:29.062] [INFO ] Processed 3000 / 3644 queries
#> [2026-07-26 22:43:38.600] [INFO ] Processed 3500 / 3644 queries
#> [2026-07-26 22:43:40.833] [INFO ] Processed 3644 / 3644 queries
#> [2026-07-26 22:43:40.865] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=872248] (1m 22s)
#> [2026-07-26 22:43:41.047] [INFO ] > Starting: harmonize_adducts [n_rows=549394]
#> [2026-07-26 22:43:41.087] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=6, n_unique_after=6] (40ms)
#> [2026-07-26 22:43:48.691] [INFO ] > Starting: harmonize_adducts [n_rows=272264]
#> [2026-07-26 22:43:48.836] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=83, n_unique_after=80] (144ms)
#> [2026-07-26 22:43:49.321] [INFO ]
#> library spectra unique_structures Pct_spectra
#> gnps 124874 11514 100.00%
#> [2026-07-26 22:43:49.323] [INFO ] > Starting: calculate_entropy_similarity [n_library=124874, n_query=3660, method=gnps]
#> [2026-07-26 22:43:49.324] [INFO ] Calculating entropy and similarity for 3660 spectra
#> [2026-07-26 22:43:52.325] [INFO ] Processed 500 / 3660 queries
#> [2026-07-26 22:43:55.052] [INFO ] Processed 1000 / 3660 queries
#> [2026-07-26 22:43:57.456] [INFO ] Processed 1500 / 3660 queries
#> [2026-07-26 22:43:59.018] [INFO ] Processed 2000 / 3660 queries
#> [2026-07-26 22:44:01.236] [INFO ] Processed 2500 / 3660 queries
#> [2026-07-26 22:44:02.261] [INFO ] Processed 3000 / 3660 queries
#> [2026-07-26 22:44:03.901] [INFO ] Processed 3500 / 3660 queries
#> [2026-07-26 22:44:04.146] [INFO ] Processed 3660 / 3660 queries
#> [2026-07-26 22:44:04.170] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=177437] (14.8s)
#> [2026-07-26 22:44:04.175] [INFO ] > Starting: harmonize_adducts [n_rows=124874]
#> [2026-07-26 22:44:04.186] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=51, n_unique_after=49] (11ms)
#> [2026-07-26 22:44:04.349] [INFO ] > Starting: calculate_entropy_similarity [n_library=94530, n_query=3644, method=gnps]
#> [2026-07-26 22:44:04.355] [INFO ] Calculating entropy and similarity for 3644 spectra
#> [2026-07-26 22:44:07.562] [INFO ] Processed 500 / 3644 queries
#> [2026-07-26 22:44:10.125] [INFO ] Processed 1000 / 3644 queries
#> [2026-07-26 22:44:12.814] [INFO ] Processed 1500 / 3644 queries
#> [2026-07-26 22:44:15.161] [INFO ] Processed 2000 / 3644 queries
#> [2026-07-26 22:44:17.176] [INFO ] Processed 2500 / 3644 queries
#> [2026-07-26 22:44:18.262] [INFO ] Processed 3000 / 3644 queries
#> [2026-07-26 22:44:20.100] [INFO ] Processed 3500 / 3644 queries
#> [2026-07-26 22:44:20.409] [INFO ] Processed 3644 / 3644 queries
#> [2026-07-26 22:44:20.434] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=212180] (16.1s)
#> [2026-07-26 22:44:20.472] [INFO ] > Starting: harmonize_adducts [n_rows=124870]
#> [2026-07-26 22:44:20.484] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=51, n_unique_after=49] (12ms)
#> [2026-07-26 22:44:23.887] [INFO ] > Starting: harmonize_adducts [n_rows=62855]
#> [2026-07-26 22:44:23.947] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=64, n_unique_after=62] (60ms)
#> [2026-07-26 22:44:24.153] [INFO ]
#> library spectra unique_structures Pct_spectra
#> massbank 29650 3310 100.00%
#> [2026-07-26 22:44:24.155] [INFO ] > Starting: calculate_entropy_similarity [n_library=29650, n_query=3660, method=gnps]
#> [2026-07-26 22:44:24.156] [INFO ] Calculating entropy and similarity for 3660 spectra
#> [2026-07-26 22:44:24.814] [INFO ] Processed 500 / 3660 queries
#> [2026-07-26 22:44:25.319] [INFO ] Processed 1000 / 3660 queries
#> [2026-07-26 22:44:25.768] [INFO ] Processed 1500 / 3660 queries
#> [2026-07-26 22:44:26.735] [INFO ] Processed 2000 / 3660 queries
#> [2026-07-26 22:44:27.046] [INFO ] Processed 2500 / 3660 queries
#> [2026-07-26 22:44:27.290] [INFO ] Processed 3000 / 3660 queries
#> [2026-07-26 22:44:27.541] [INFO ] Processed 3500 / 3660 queries
#> [2026-07-26 22:44:27.596] [INFO ] Processed 3660 / 3660 queries
#> [2026-07-26 22:44:27.610] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=40342] (3.5s)
#> [2026-07-26 22:44:27.613] [INFO ] > Starting: harmonize_adducts [n_rows=29650]
#> [2026-07-26 22:44:27.618] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=48, n_unique_after=46] (5ms)
#> [2026-07-26 22:44:27.749] [INFO ] > Starting: calculate_entropy_similarity [n_library=21301, n_query=3644, method=gnps]
#> [2026-07-26 22:44:27.752] [INFO ] Calculating entropy and similarity for 3644 spectra
#> [2026-07-26 22:44:28.372] [INFO ] Processed 500 / 3644 queries
#> [2026-07-26 22:44:29.477] [INFO ] Processed 1000 / 3644 queries
#> [2026-07-26 22:44:29.874] [INFO ] Processed 1500 / 3644 queries
#> [2026-07-26 22:44:30.215] [INFO ] Processed 2000 / 3644 queries
#> [2026-07-26 22:44:30.524] [INFO ] Processed 2500 / 3644 queries
#> [2026-07-26 22:44:30.815] [INFO ] Processed 3000 / 3644 queries
#> [2026-07-26 22:44:31.077] [INFO ] Processed 3500 / 3644 queries
#> [2026-07-26 22:44:31.152] [INFO ] Processed 3644 / 3644 queries
#> [2026-07-26 22:44:31.167] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=38688] (3.4s)
#> [2026-07-26 22:44:31.182] [INFO ] > Starting: harmonize_adducts [n_rows=29649]
#> [2026-07-26 22:44:31.187] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=48, n_unique_after=46] (5ms)
#> [2026-07-26 22:44:37.299] [INFO ] > Starting: harmonize_adducts [n_rows=336677]
#> [2026-07-26 22:44:37.324] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=7, n_unique_after=7] (24ms)
#> [2026-07-26 22:44:37.910] [INFO ]
#> library spectra unique_structures Pct_spectra
#> merlin 203603 26316 100.00%
#> [2026-07-26 22:44:37.911] [INFO ] > Starting: calculate_entropy_similarity [n_library=203603, n_query=3660, method=gnps]
#> [2026-07-26 22:44:37.912] [INFO ] Calculating entropy and similarity for 3660 spectra
#> [2026-07-26 22:44:43.208] [INFO ] Processed 500 / 3660 queries
#> [2026-07-26 22:44:46.962] [INFO ] Processed 1000 / 3660 queries
#> [2026-07-26 22:44:50.885] [INFO ] Processed 1500 / 3660 queries
#> [2026-07-26 22:44:53.989] [INFO ] Processed 2000 / 3660 queries
#> [2026-07-26 22:44:56.641] [INFO ] Processed 2500 / 3660 queries
#> [2026-07-26 22:44:59.060] [INFO ] Processed 3000 / 3660 queries
#> [2026-07-26 22:45:01.266] [INFO ] Processed 3500 / 3660 queries
#> [2026-07-26 22:45:01.724] [INFO ] Processed 3660 / 3660 queries
#> [2026-07-26 22:45:01.752] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=282510] (23.8s)
#> [2026-07-26 22:45:01.758] [INFO ] > Starting: harmonize_adducts [n_rows=203603]
#> [2026-07-26 22:45:01.774] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=7, n_unique_after=7] (16ms)
#> [2026-07-26 22:45:01.938] [INFO ] > Starting: calculate_entropy_similarity [n_library=142636, n_query=3644, method=gnps]
#> [2026-07-26 22:45:01.948] [INFO ] Calculating entropy and similarity for 3644 spectra
#> [2026-07-26 22:45:06.575] [INFO ] Processed 500 / 3644 queries
#> [2026-07-26 22:45:10.563] [INFO ] Processed 1000 / 3644 queries
#> [2026-07-26 22:45:14.559] [INFO ] Processed 1500 / 3644 queries
#> [2026-07-26 22:45:17.751] [INFO ] Processed 2000 / 3644 queries
#> [2026-07-26 22:45:20.590] [INFO ] Processed 2500 / 3644 queries
#> [2026-07-26 22:45:23.009] [INFO ] Processed 3000 / 3644 queries
#> [2026-07-26 22:45:25.289] [INFO ] Processed 3500 / 3644 queries
#> [2026-07-26 22:45:25.738] [INFO ] Processed 3644 / 3644 queries
#> [2026-07-26 22:45:25.763] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=283632] (23.8s)
#> [2026-07-26 22:45:25.833] [INFO ] > Starting: harmonize_adducts [n_rows=203519]
#> [2026-07-26 22:45:25.851] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=7, n_unique_after=7] (18ms)
#> [2026-07-26 22:45:29.019] [INFO ] > Starting: harmonize_adducts [n_rows=32610]
#> [2026-07-26 22:45:29.193] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=138, n_unique_after=136] (174ms)
#> [2026-07-26 22:45:29.403] [INFO ]
#> library spectra unique_structures Pct_spectra
#> multims2 19091 1481 100.00%
#> [2026-07-26 22:45:29.405] [INFO ] > Starting: calculate_entropy_similarity [n_library=19091, n_query=3660, method=gnps]
#> [2026-07-26 22:45:29.406] [INFO ] Calculating entropy and similarity for 3660 spectra
#> [2026-07-26 22:45:29.810] [INFO ] Processed 500 / 3660 queries
#> [2026-07-26 22:45:30.084] [INFO ] Processed 1000 / 3660 queries
#> [2026-07-26 22:45:30.389] [INFO ] Processed 1500 / 3660 queries
#> [2026-07-26 22:45:30.615] [INFO ] Processed 2000 / 3660 queries
#> [2026-07-26 22:45:30.818] [INFO ] Processed 2500 / 3660 queries
#> [2026-07-26 22:45:30.997] [INFO ] Processed 3000 / 3660 queries
#> [2026-07-26 22:45:31.212] [INFO ] Processed 3500 / 3660 queries
#> [2026-07-26 22:45:31.843] [INFO ] Processed 3660 / 3660 queries
#> [2026-07-26 22:45:31.854] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=25448] (2.4s)
#> [2026-07-26 22:45:31.858] [INFO ] > Starting: harmonize_adducts [n_rows=19091]
#> [2026-07-26 22:45:31.862] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=105, n_unique_after=104] (4ms)
#> [2026-07-26 22:45:32.010] [INFO ] > Starting: calculate_entropy_similarity [n_library=12761, n_query=3644, method=gnps]
#> [2026-07-26 22:45:32.012] [INFO ] Calculating entropy and similarity for 3644 spectra
#> [2026-07-26 22:45:32.360] [INFO ] Processed 500 / 3644 queries
#> [2026-07-26 22:45:32.687] [INFO ] Processed 1000 / 3644 queries
#> [2026-07-26 22:45:32.980] [INFO ] Processed 1500 / 3644 queries
#> [2026-07-26 22:45:33.195] [INFO ] Processed 2000 / 3644 queries
#> [2026-07-26 22:45:33.370] [INFO ] Processed 2500 / 3644 queries
#> [2026-07-26 22:45:33.542] [INFO ] Processed 3000 / 3644 queries
#> [2026-07-26 22:45:33.668] [INFO ] Processed 3500 / 3644 queries
#> [2026-07-26 22:45:33.707] [INFO ] Processed 3644 / 3644 queries
#> [2026-07-26 22:45:33.716] [INFO ] [OK] Completed: calculate_entropy_similarity [n_comparisons=21750] (1.7s)
#> [2026-07-26 22:45:33.726] [INFO ] > Starting: harmonize_adducts [n_rows=19076]
#> [2026-07-26 22:45:33.731] [INFO ] [OK] Completed: harmonize_adducts [n_unique_before=105, n_unique_after=104] (5ms)
#> [2026-07-26 22:45:36.415] [INFO ] Annotating 3660 query spectra against 1450127 library references
#> [2026-07-26 22:45:53.223] [INFO ] Here is the distribution of annotation similarity scores (0.1 bins):
#> [2026-07-26 22:45:53.226] [INFO ]
#> bin N Pct
#> [0,0.1] 1342565 88.06%
#> (0.1,0.2] 109588 7.19%
#> (0.2,0.3] 38101 2.50%
#> (0.3,0.4] 16637 1.09%
#> (0.4,0.5] 8610 0.56%
#> (0.5,0.6] 4604 0.30%
#> (0.6,0.7] 1627 0.11%
#> (0.7,0.8] 1063 0.07%
#> (0.8,0.9] 1260 0.08%
#> (0.9,1] 516 0.03%
#> [2026-07-26 22:45:53.319] [INFO ] 591373 Candidates annotated on 3639 features (threshold >= 0).
#> [2026-07-26 22:45:53.323] [INFO ] Exporting parameters to: data/interim/params/260726_224553_annotate_spectra.yaml
#> [2026-07-26 22:45:53.326] [INFO ] > Starting: export_output [file=data/interim/annotations/example_spectralMatches_pos.tsv.gz, n_rows=1524571]
#> [2026-07-26 22:45:58.727] [INFO ] [OK] Completed: export_output [size_bytes=88043670] (5.4s)
#> ✔ ann_spe_pos completed [8m 18.5s, 88.04 MB]
#> + ann_spe_neg dispatched
#> [2026-07-26 22:46:01.456] [INFO ] ============================================================
#> [2026-07-26 22:46:01.458] [INFO ] Data Sanitizing: Pre-flight Checks
#> [2026-07-26 22:46:01.459] [INFO ] ============================================================
#> [2026-07-26 22:46:01.461] [INFO ] Checking MGF file...
#> [2026-07-26 22:46:01.993] [INFO ] [OK] MGF file: 12195 MS2 spectra found
#> [2026-07-26 22:46:01.995] [INFO ] ============================================================
#> [2026-07-26 22:46:01.996] [INFO ] [OK] All pre-flight checks passed!
#> [2026-07-26 22:46:01.997] [INFO ] Data validation complete. Ready to proceed.
#> [2026-07-26 22:46:01.997] [INFO ] ============================================================
#> [2026-07-26 22:46:01.998] [INFO ] Starting spectral annotation in neg mode
#> [2026-07-26 22:46:01.999] [INFO ] Importing spectra from: data/source/example_spectra.mgf
#> [2026-07-26 22:46:02.001] [INFO ] Reading MGF file (7.41 MB) with optimized parser: data/source/example_spectra.mgf
#> [2026-07-26 22:46:04.508] [INFO ] Processed 10000 spectra...
#> [2026-07-26 22:46:05.758] [INFO ] Total spectra read: 16282
#> [2026-07-26 22:46:13.600] [INFO ] Loaded 16282 spectra from file
#> [2026-07-26 22:46:13.616] [INFO ] Combining replicate spectra by FEATURE_ID
#> [2026-07-26 22:46:13.620] [INFO ] Combined replicates: 0 -> 0 spectra
#> [2026-07-26 22:46:13.661] [WARN ] No spectra to sanitize
#> [2026-07-26 22:46:13.662] [INFO ] Import complete: 0 spectra ready for analysis
#> [2026-07-26 22:46:13.663] [WARN ] No query spectra loaded
#> [2026-07-26 22:46:13.666] [INFO ] Exporting parameters to: data/interim/params/260726_224613_annotate_spectra.yaml
#> [2026-07-26 22:46:13.668] [WARN ] Returning empty annotation template
#> [2026-07-26 22:46:13.671] [INFO ] > Starting: export_output [file=data/interim/annotations/example_spectralMatches_neg.tsv.gz, n_rows=1]
#> [2026-07-26 22:46:13.673] [INFO ] [OK] Completed: export_output [size_bytes=308] (2ms)
#> ✔ ann_spe_neg completed [12.3s, 308 B]
#> + fea_edg_pre dispatched
#> [2026-07-26 22:46:15.972] [INFO ] > Starting: prepare_features_edges [n_edge_types=2]
#> [2026-07-26 22:46:16.064] [INFO ] [OK] Completed: prepare_features_edges [n_edges=49386] (92ms)
#> [2026-07-26 22:46:16.088] [INFO ] Exporting parameters to: data/interim/params/260726_224616_prepare_features_edges.yaml
#> [2026-07-26 22:46:16.089] [INFO ] > Starting: export_output [file=data/interim/features/example_edges.tsv, n_rows=49386]
#> [2026-07-26 22:46:16.097] [INFO ] [OK] Completed: export_output [size_bytes=2574921] (7ms)
#> ✔ fea_edg_pre completed [134ms, 2.57 MB]
#> + ann_spe_pre dispatched
#> [2026-07-26 22:46:18.292] [INFO ] Preparing spectral matching annotations from 2 file(s)
#> [2026-07-26 22:46:36.899] [INFO ] > Starting: process_smiles [n_structures=591373]
#> [2026-07-26 22:46:36.901] [INFO ] Processing SMILES with RDKit
#> [2026-07-26 22:46:50.794] [INFO ] Processing 40 new SMILES with RDKit
#> [2026-07-26 22:46:50.795] [INFO ] Starting SMILES processing pipeline
#> [2026-07-26 22:46:50.795] [INFO ] Input: /tmp/RtmpZ8Qf0D/file26e550b91a08.smi
#> [2026-07-26 22:46:50.795] [INFO ] Output: /tmp/RtmpZ8Qf0D/file26e576401495.csv.gz
#> [2026-07-26 22:46:50.796] [INFO ] Input file validated: /tmp/RtmpZ8Qf0D/file26e550b91a08.smi
#> [2026-07-26 22:46:50.796] [INFO ] Output file validated: /tmp/RtmpZ8Qf0D/file26e576401495.csv.gz
#> [2026-07-26 22:46:50.796] [INFO ] Processing parameters: workers=8, batch_size=1000, progress_interval=10000
#> [2026-07-26 22:46:50.796] [INFO ] SMILES supplier initialized
#> [2026-07-26 22:46:50.860] [INFO ] Processing complete. Total molecules processed: 40
#> [2026-07-26 22:46:50.901] [INFO ] Successfully processed 40 SMILES
#> [2026-07-26 22:46:59.450] [INFO ] [OK] Completed: process_smiles [n_processed=591373] (22.6s)
#> [2026-07-26 22:47:13.731] [INFO ] > Starting: complement_metadata [n_input=1524571]
#> [2026-07-26 22:47:52.933] [INFO ] [OK] Completed: complement_metadata [n_enriched=1524571] (39.2s)
#> [2026-07-26 22:47:53.409] [INFO ] Exporting parameters to: data/interim/params/260726_224753_prepare_annotations_spectra.yaml
#> [2026-07-26 22:47:53.411] [INFO ] > Starting: export_output [file=data/interim/annotations/example_spectralMatchesPrepared.tsv.gz, n_rows=1524571]
#> [2026-07-26 22:48:01.516] [INFO ] [OK] Completed: export_output [size_bytes=153539673] (8.1s)
#> ✔ ann_spe_pre completed [1m 43.2s, 153.54 MB]
#> + fea_com dispatched
#> [2026-07-26 22:48:04.251] [INFO ] > Starting: create_components [n_input_files=1]
#> [2026-07-26 22:48:04.253] [INFO ] Creating components from 1 edge file(s)
#> [2026-07-26 22:48:04.310] [INFO ] Loaded 46222 edges connecting 4584 unique features
#> [2026-07-26 22:48:04.694] [INFO ] Selected community partition at resolution 0.025 (modularity 0.805)
#> [2026-07-26 22:48:04.714] [INFO ] Found 803 communities using weighted Louvain/Leiden clustering
#> [2026-07-26 22:48:04.717] [INFO ] Component sizes - Min: 1, Max: 360, Mean: 5.7
#> [2026-07-26 22:48:04.735] [INFO ] Exporting parameters to: data/interim/params/260726_224804_create_components.yaml
#> [2026-07-26 22:48:04.737] [INFO ] > Starting: export_output [file=data/interim/features/example_components.tsv, n_rows=4584]
#> [2026-07-26 22:48:04.739] [INFO ] [OK] Completed: export_output [size_bytes=40698] (2ms)
#> [2026-07-26 22:48:04.740] [INFO ] Components written to: data/interim/features/example_components.tsv
#> [2026-07-26 22:48:04.741] [INFO ] [OK] Completed: create_components [n_components=803, n_features=4584] (490ms)
#> ✔ fea_com completed [497ms, 40.70 kB]
#> + ann_fil dispatched
#> [2026-07-26 22:48:07.001] [INFO ] > Starting: filter_annotations [n_annotation_files=6, tolerance_rt=Inf]
#> [2026-07-26 22:48:07.003] [INFO ] Filtering annotations
#> [2026-07-26 22:48:07.060] [INFO ] Processing 5328 unique features for annotation filtering
#> [2026-07-26 22:48:18.714] [INFO ] Removing MS1 annotations superseded by quality spectral matches
#> [2026-07-26 22:48:22.022] [INFO ] Removed 16331 redundant MS1 annotations
#> [2026-07-26 22:48:22.023] [INFO ] Total annotations after MS1 deduplication: 2121587
#> [2026-07-26 22:48:41.721] [INFO ] Adduct-semantics filter: before=2121612, removed_total=0, removed_spectral_mismatch=0, after=2121612
#> [2026-07-26 22:48:41.761] [INFO ] Joining RT library and computing RT deltas
#> [2026-07-26 22:48:50.408] [INFO ] Removed 10201 duplicate RT library matches (keeping best match per annotation)
#> [2026-07-26 22:48:50.414] [INFO ] RT deltas computed for 0 annotations (no hard cutoff applied; scoring handles RT penalty)
#> [2026-07-26 22:48:50.417] [INFO ] Removed 10201 duplicate RT library matches during join
#> [2026-07-26 22:48:51.206] [INFO ] Exporting parameters to: data/interim/params/260726_224851_filter_annotations.yaml
#> [2026-07-26 22:48:51.209] [INFO ] > Starting: export_output [file=data/interim/annotations/example_annotationsFiltered.tsv.gz, n_rows=2111411]
#> [2026-07-26 22:49:02.176] [INFO ] [OK] Completed: export_output [size_bytes=198256644] (11s)
#> [2026-07-26 22:49:02.178] [INFO ] [OK] Completed: filter_annotations [n_filtered=2111411] (55.2s)
#> ✔ ann_fil completed [55.2s, 198.26 MB]
#> + fea_com_pre dispatched
#> [2026-07-26 22:49:05.069] [INFO ] > Starting: prepare_features_components [n_files=1]
#> [2026-07-26 22:49:05.077] [INFO ] [OK] Completed: prepare_features_components [n_assignments=4584] (8ms)
#> [2026-07-26 22:49:05.097] [INFO ] Exporting parameters to: data/interim/params/260726_224905_prepare_features_components.yaml
#> [2026-07-26 22:49:05.099] [INFO ] > Starting: export_output [file=data/interim/features/example_componentsPrepared.tsv, n_rows=4584]
#> [2026-07-26 22:49:05.101] [INFO ] [OK] Completed: export_output [size_bytes=40693] (2ms)
#> ✔ fea_com_pre completed [38ms, 40.69 kB]
#> + ann_wei dispatched
#> [2026-07-26 22:49:07.267] [INFO ] Starting annotation weighting and scoring
#> [2026-07-26 22:49:07.268] [INFO ] > Starting: weight_annotations [n_candidates_neighbors=16, n_candidates_final=1]
#> [2026-07-26 22:49:50.212] [INFO ]
#> candidate_library n Pct
#> ISDB - Wikidata 1015532 48.10%
#> TIMA MS1 584830 27.70%
#> enveda180 322220 15.26%
#> ISDB - NormanSusDat 73712 3.49%
#> merlin 66921 3.17%
#> gnps 33873 1.60%
#> massbank 8197 0.39%
#> multims2 3405 0.16%
#> SIRIUS 2561 0.12%
#> [2026-07-26 22:50:10.561] [INFO ] > Starting: weight_bio [n_annotations=1834685, n_sop=1334047]
#> [2026-07-26 22:50:10.565] [INFO ] Weighting 1834685 annotations by biological source
#> [2026-07-26 22:50:22.713] [INFO ] [OK] Completed: weight_bio [n_weighted=1834685] (12.2s)
#> [2026-07-26 22:50:22.715] [INFO ] > Starting: decorate_bio [n_annotations=1834685]
#> [2026-07-26 22:50:24.415] [INFO ] Taxonomically informed metabolite annotation reranked:
#> Kingdom level: 243682 candidates (59840 unique)
#> Phylum level: 241276 candidates (59108 unique)
#> Class level: 192351 candidates (51215 unique)
#> Order level: 40222 candidates (12756 unique)
#> Family level: 32435 candidates (10062 unique)
#> Tribe level: 6121 candidates (1544 unique)
#> Genus level: 5078 candidates (1203 unique)
#> Species level: 2865 candidates (610 unique)
#> Variety level: 522 candidates (133 unique)
#> Biota level: 522 candidates (133 unique)
#> [2026-07-26 22:50:24.416] [INFO ] [OK] Completed: decorate_bio [n_processed=1834685] (1.7s)
#> [2026-07-26 22:50:24.418] [INFO ] > Starting: clean_bio [n_annotations=1834685, minimal_consistency=0]
#> [2026-07-26 22:51:05.772] [INFO ] [OK] Completed: clean_bio [n_cleaned=1834685] (41.4s)
#> [2026-07-26 22:51:05.777] [INFO ] > Starting: weight_chemo [n_input=1834685]
#> [2026-07-26 22:51:05.779] [INFO ] Weighting 1834685 annotations by chemical consistency
#> [2026-07-26 22:51:11.930] [INFO ] [OK] Completed: weight_chemo [n_weighted=1834685] (6.2s)
#> [2026-07-26 22:51:11.933] [INFO ] > Starting: combine_weighted_scores [n_bio=1834685, n_chemo=1834685]
#> [2026-07-26 22:51:13.465] [INFO ] [OK] Completed: combine_weighted_scores [n_combined=1834685] (1.5s)
#> [2026-07-26 22:51:13.467] [INFO ] > Starting: decorate_chemo [n_annotations=1834685]
#> [2026-07-26 22:51:25.082] [INFO ] Chemically informed metabolite annotation reranked:
#> Classyfire:
#> Kingdom level: 725606 candidates (332816 unique)
#> Superclass level: 518393 candidates (227214 unique)
#> Class level: 442889 candidates (191172 unique)
#> Parent level: 393425 candidates (170448 unique)
#> NPClassifier:
#> Pathway level: 730687 candidates (335741 unique)
#> Superclass level: 511768 candidates (224429 unique)
#> Class level: 396782 candidates (172256 unique)
#> [2026-07-26 22:51:25.084] [INFO ] [OK] Completed: decorate_chemo [n_processed=1834685] (11.6s)
#> [2026-07-26 22:51:35.623] [INFO ] > Starting: clean_chemo [n_annotations=1834685, candidates_final=1, high_evidence=FALSE]
#> [2026-07-26 22:52:21.989] [INFO ] Enforced cluster entity consensus for 4482 features (anchor InChIKey promoted to rank 1)
#> [2026-07-26 22:53:09.174] [INFO ] Sampling candidates for 644 features with more than 7 candidates per score
#> [2026-07-26 22:53:17.569] [INFO ] > Starting: filter_high_evidence [n_input=1822127, context=filtered]
#> [2026-07-26 22:53:19.938] [INFO ] [filtered] Removed 1819564 low-evidence candidates (99.9% of 1822127 total)
#> [2026-07-26 22:53:19.939] [INFO ] [filtered] 2563 high-evidence candidates remaining (0.1%)
#> [2026-07-26 22:53:19.940] [INFO ] [OK] Completed: filter_high_evidence [n_filtered=2563, n_removed=1819564] (2.4s)
#> [2026-07-26 22:53:20.302] [INFO ] Filtered-tier promotion gate removed 58 non-rank1 row(s) for promoted features
#> [2026-07-26 22:53:20.304] [INFO ] Summarizing annotation results
#> [2026-07-26 22:53:21.489] [INFO ] Annotated features: 911/5328 (17.1%)
#> [2026-07-26 22:53:28.509] [INFO ] Summarizing annotation results
#> [2026-07-26 22:54:07.357] [INFO ] Annotated features: 5231/5328 (98.2%)
#> [2026-07-26 22:54:10.993] [INFO ] [OK] Completed: clean_chemo [n_final_full=1822127, n_final_filtered=911, n_final_mini=5328, n_features=911] (2m 35s)
#> [2026-07-26 22:54:10.995] [INFO ] [OK] Completed: weight_annotations [n_annotations=NULL] (5m 4s)
#> [2026-07-26 22:54:11.025] [INFO ] Exporting parameters to: data/processed/20260726_225410_example/260726_225411_prepare_params.yaml
#> [2026-07-26 22:54:11.050] [INFO ] Exporting parameters to: data/processed/20260726_225410_example/260726_225411_prepare_params_advanced.yaml
#> [2026-07-26 22:54:11.053] [INFO ] > Starting: export_output [file=data/processed/20260726_225410_example/example_results_mini.tsv, n_rows=5328]
#> [2026-07-26 22:54:11.058] [INFO ] [OK] Completed: export_output [size_bytes=912244] (4ms)
#> [2026-07-26 22:54:11.060] [INFO ] > Starting: export_output [file=data/processed/20260726_225410_example/example_results_filtered.tsv, n_rows=911]
#> [2026-07-26 22:54:11.064] [INFO ] [OK] Completed: export_output [size_bytes=684377] (4ms)
#> [2026-07-26 22:54:11.066] [INFO ] > Starting: export_output [file=data/processed/20260726_225410_example/example_results.tsv, n_rows=1822127]
#> [2026-07-26 22:54:13.998] [INFO ] [OK] Completed: export_output [size_bytes=866245095] (2.9s)
#> [2026-07-26 22:54:14.000] [INFO ] Results exported: example_results.tsv
#> ✔ ann_wei completed [5m 6.7s, 866.93 MB]
#> + exp_mzt dispatched
#> [2026-07-26 22:54:19.428] [INFO ] > Starting: write_mztab [input=example_results_filtered.tsv, output=example_results.mztab]
#> [2026-07-26 22:54:19.703] [INFO ] [OK] Completed: write_mztab [n_sml=911, n_smf=911, n_sme=911] (275ms)
#> ✔ exp_mzt completed [346ms, 1.08 MB]
#> ✔ ended pipeline [31m 58s, 157 completed, 0 skipped]
#> There were 13 warnings (use warnings() to see them)3 Performing Taxonomically Informed Metabolite Annotation
This vignette describes how Taxonomically Informed Metabolite Annotation is performed. If you followed all previous steps successfully, this should be a piece of cake, you deserve it!
The final exported file is formatted in order to be easily imported in Cytoscape to further explore your data!
We hope you enjoyed using TIMA and are pleased to hear from you!
For any remark or suggestion, please fill an issue or feel free to contact us directly.
Reuse
Citation
BibTeX citation:
@online{rutz2026,
author = {Rutz, Adriano},
title = {3 {Performing} {Taxonomically} {Informed} {Metabolite}
{Annotation}},
date = {2026-07-26},
url = {https://taxonomicallyinformedannotation.github.io/tima/vignettes/articles/III-processing.html},
langid = {en}
}
For attribution, please cite this work as:
Rutz, Adriano. 2026. “3 Performing Taxonomically Informed
Metabolite Annotation.” July 26. https://taxonomicallyinformedannotation.github.io/tima/vignettes/articles/III-processing.html.