{"id":113274,"date":"2026-07-29T07:25:44","date_gmt":"2026-07-29T14:25:44","guid":{"rendered":"https:\/\/www.backblaze.com\/blog\/?p=113274"},"modified":"2026-07-29T07:25:46","modified_gmt":"2026-07-29T14:25:46","slug":"ai-data-pipeline-101-ingest-archive","status":"publish","type":"post","link":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/","title":{"rendered":"AI Data Pipeline 101: Ingest &amp; Archive"},"content":{"rendered":"\n<figure class=\"wp-block-image size-large\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"583\" src=\"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/AI-0002-Blog-Header-1440x820-1-1024x583.png\" alt=\"An illustration of gears, boxes, and a graphic that says AI.\" class=\"wp-image-113277\" srcset=\"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/AI-0002-Blog-Header-1440x820-1-1024x583.png 1024w, https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/AI-0002-Blog-Header-1440x820-1-300x171.png 300w, https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/AI-0002-Blog-Header-1440x820-1-768x437.png 768w, https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/AI-0002-Blog-Header-1440x820-1.png 1440w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><\/figure>\n\n\n\n<div style=\"height:15px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<p class=\"wp-block-paragraph\">Extensive news coverage and analyst reports on AI missing productivity and ROI targets mean that AI failure is something of a hot topic. There\u2019s no arguing that some AI initiatives are misguided, including replacing entire specialist teams with AI. For others, the issue actually lies with data knowledge and readiness\u2013<a href=\"https:\/\/www.gartner.com\/en\/newsroom\/press-releases\/2025-02-26-lack-of-ai-ready-data-puts-ai-projects-at-risk\" target=\"_blank\" rel=\"noreferrer noopener\">Gartner predicts<\/a> that through 2026, organizations will abandon 60% of AI projects unsupported by AI-ready data, and <a href=\"https:\/\/www.spglobal.com\/market-intelligence\/en\/news-insights\/videos\/unlocking-ai-driven-insights-for-risk-management\" target=\"_blank\" rel=\"noreferrer noopener\">S&amp;P Global recently highlighted<\/a> the importance of ingesting previously overlooked or unknown data to discover interdependencies in risk management.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">The pressure to move at the perceived speed of AI makes it easy to skip or rush important prep work. Now that major AI pilots have been up and running, supplementing or sometimes entirely replacing select business functions <a href=\"https:\/\/www.spglobal.com\/market-intelligence\/en\/news-insights\/research\/2025\/10\/generative-ai-shows-rapid-growth-but-yields-mixed-results\" target=\"_blank\" rel=\"noreferrer noopener\">with mixed results<\/a>, this exposes an already known problem among AI experts: implementing AI too quickly and ignoring the importance of keeping human experts in the loop increases your threshold for error. This is especially prevalent for both internal AI tools that are meant to augment key roles, and customer-facing AI applications that are supposed to increase accessibility to an outcome, such as generating a lifelike video based on natural language prompts.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">What\u2019s causing this? It\u2019s not just the LLMs\u2013it\u2019s the data. Now that human experts are more aware of what AI can get wrong, we\u2019re going back to basics to help you get it right. This starts at the very beginning: curating, ingesting, and storing data using infrastructure that\u2019s actually designed for moving data quickly, and without financial penalty.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">The Stages of the AI Model Training Data Pipeline<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Our in-house AI experts have <a href=\"https:\/\/www.backblaze.com\/blog\/architecting-your-ai-data-pipeline-using-b2-overdrive\/\" target=\"_blank\" rel=\"noreferrer noopener\">split the AI data pipeline into five essential key phases<\/a> for model training, and highlighting how storage factors into each stage (including what\u2019s being stored.)<\/p>\n\n\n\n<figure data-wp-context=\"{&quot;imageId&quot;:&quot;6a6a2fade64d4&quot;}\" data-wp-interactive=\"core\/image\" data-wp-key=\"6a6a2fade64d4\" class=\"wp-block-image size-large wp-lightbox-container\"><img loading=\"lazy\" decoding=\"async\" width=\"1024\" height=\"401\" data-wp-class--hide=\"state.isContentHidden\" data-wp-class--show=\"state.isContentVisible\" data-wp-init=\"callbacks.setButtonStyles\" data-wp-on--click=\"actions.showLightbox\" data-wp-on--load=\"callbacks.setButtonStyles\" data-wp-on--pointerdown=\"actions.preloadImage\" data-wp-on--pointerenter=\"actions.preloadImageWithDelay\" data-wp-on--pointerleave=\"actions.cancelPreload\" data-wp-on-window--resize=\"callbacks.setButtonStyles\" src=\"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/image1-1024x401.png\" alt=\"An infographic of the AI data pipeline.\" class=\"wp-image-113276\" srcset=\"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/image1-1024x401.png 1024w, https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/image1-300x118.png 300w, https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/image1-768x301.png 768w, https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/image1-1536x602.png 1536w, https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/image1-1568x614.png 1568w, https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/image1.png 1868w\" sizes=\"auto, (max-width: 1024px) 100vw, 1024px\" \/><button\n\t\t\tclass=\"lightbox-trigger\"\n\t\t\ttype=\"button\"\n\t\t\taria-haspopup=\"dialog\"\n\t\t\tdata-wp-bind--aria-label=\"state.thisImage.triggerButtonAriaLabel\"\n\t\t\tdata-wp-init=\"callbacks.initTriggerButton\"\n\t\t\tdata-wp-on--click=\"actions.showLightbox\"\n\t\t\tdata-wp-style--right=\"state.thisImage.buttonRight\"\n\t\t\tdata-wp-style--top=\"state.thisImage.buttonTop\"\n\t\t>\n\t\t\t<svg xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"12\" height=\"12\" fill=\"none\" viewBox=\"0 0 12 12\">\n\t\t\t\t<path fill=\"#fff\" d=\"M2 0a2 2 0 0 0-2 2v2h1.5V2a.5.5 0 0 1 .5-.5h2V0H2Zm2 10.5H2a.5.5 0 0 1-.5-.5V8H0v2a2 2 0 0 0 2 2h2v-1.5ZM8 12v-1.5h2a.5.5 0 0 0 .5-.5V8H12v2a2 2 0 0 1-2 2H8Zm2-12a2 2 0 0 1 2 2v2h-1.5V2a.5.5 0 0 0-.5-.5H8V0h2Z\" \/>\n\t\t\t<\/svg>\n\t\t<\/button><\/figure>\n\n\n\n<div style=\"height:15px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<h2 class=\"wp-block-heading\">Data collection \u2260 data ingest<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">Data ingest is the process of any type of data being added to a designated collection destination, whether that\u2019s a specific file folder, database, or object storage bucket. For a lot of applications, data ingest is frequent or nearly constant\u2013busy e-commerce sites with a constant flow of customer transactions and feedback, live video feeds, and combining real-time data sources like pairing security footage with physical building security sensors. This is called streaming ingestion\u2013and streaming ingestion being the foundation of data collection for various types of AI is one of the key reasons why storage is an AI infrastructure problem that often flies under the radar until there\u2019s a serious problem.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Streaming ingestion requires constant low-latency access to the data storage repository to prevent data upload lags and errors.<\/li>\n\n\n\n<li>Streaming ingestion for video and other large files requires high rate limits and high-throughput networking capabilities to optimize upload times, especially when an application requires a file to be ready for processing in seconds\/minutes instead of hours\/days.<\/li>\n\n\n\n<li>Running out of storage capacity is not an option for model performance, and for compliance and auditing purposes\u2013streaming ingestion requires constant access to storage that is as close to infinite as possible.<\/li>\n<\/ul>\n\n\n\n<h2 class=\"wp-block-heading\">Automation from the start<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">In addition to the files themselves, setting up a highly effective AI data pipeline involves building automation from the very beginning to immediately allocate files to the right bucket using taxonomy and collect and store file metadata to begin data aggregation critical for future labeling and processing.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Why taxonomy is critical for AI<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">The file itself <em>is <\/em>the data source. Navigating your dataset starts with implementing a taxonomy that makes your dataset highly searchable as your data grows from a few thousand for your first round of training to millions for an AI application operating in production.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Taxonomy is essentially your file storage structure, or how files are automatically \u201cnested\u201d and relationships between files are built immediately upon ingest. When you\u2019re just getting started, developing your taxonomy helps you stay organized and ready to go searching for a specific file when your coworker doesn\u2019t believe what the data is saying. When you\u2019re working in established teams, implementing a new taxonomy or showing that you understand the importance of following an established taxonomy builds a contract of understanding between you and your data engineering or ML teams. (AKA, changing taxonomy mid-project is a big undertaking with ripple effects, and should not be taken lightly.)<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">A simple example taxonomy for ingesting raw files can look like this:<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><code>\/&lt;source>\/&lt;modality>\/&lt;status>\/&lt;date>\/filename<\/code><\/p>\n\n\n\n<p class=\"wp-block-paragraph\">Taxonomy should also reflect what the data is actually <em>doing <\/em>or <em>will do, <\/em>not just the file type and data source. The goal is to make every object self-describing at write time, so downstream training pipelines can filter, version, and partition without touching the data itself.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Partition by date\/time at the prefix level so you can use time-range queries and lifecycle rules without scanning everything<\/li>\n\n\n\n<li>Include modality (video, audio, image, text) as a top-level segment so cross-modal datasets stay logically separated but co-located in the same bucket<\/li>\n\n\n\n<li>Use a UUID or content hash as the filename \u2014 never rely on source filenames, which are inconsistent and collision-prone at scale<\/li>\n\n\n\n<li>Use B2 bucket policies or object tagging rules to reject objects written to non-conforming prefixes<\/li>\n\n\n\n<li>Maintain a human-readable taxonomy manifest (taxonomy.json at bucket root) that documents each top-level prefix and its schema<\/li>\n\n\n\n<li>For video specifically, consider a separate prefix segment for resolution or codec: &#8230;\/video\/4k\/h264\/&#8230; \u2014 this pays off quickly when training jobs need to filter by input spec<\/li>\n<\/ul>\n\n\n\n<h3 class=\"wp-block-heading\">Retain all the metadata<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Write metadata as object tags and custom headers at ingest using per-object user-defined metadata at PUT, such as:<\/p>\n\n\n\n<pre class=\"wp-block-preformatted\">s3.put_object(<br \/>    Bucket=\"for-training\",<br \/>    Key=object_key,<br \/>    Body=video_bytes,<br \/>    Metadata={<br \/>        \"source-id\": \"warehouse-camera-feed-8\",<br \/>        \"capture-timestamp\": \"2026-07-08T14:32:00Z\",<br \/>        \"frame-rate\": \"30\",<br \/>        \"resolution\": \"3840x2160\",<br \/>        \"modality\": \"video\",<br \/>        \"label-status\": \"unlabeled\",<br \/>        \"ingest-pipeline-version\": \"v2.3.1\"<br \/>    }<br \/>)<\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">This way, the metadata travels with the object and is always returned on HEAD requests without a separate lookup.&nbsp;<\/p>\n\n\n\n<p class=\"wp-block-paragraph\">But to build out a rich dataset, you will need even more metadata in the form of sidecar metadata files. Retain annotations, bounding boxes, ground-truth labels, licensing info, and consent flag metadata by creating a sidecar JSON with a .meta.json suffix. For the warehouse camera video example listed above, you would end up with two files that look like this:<\/p>\n\n\n\n<pre class=\"wp-block-preformatted\">for-training\/video\/warehouse-camera-feed-8\/2026\/07\/08\/14\/a3f9c1d2.mp4<br \/><br \/>for-training\/video\/warehouse-camera-feed-8\/2026\/07\/08\/14\/a3f9c1d2.meta.json<\/pre>\n\n\n\n<p class=\"wp-block-paragraph\">Enable object versioning on training buckets. If a labeling pipeline updates an annotation, it should write a new version (or a new sidecar) rather than overwriting \u2014 training reproducibility depends on knowing which metadata was present at the time a dataset was compiled.<\/p>\n\n\n\n<h3 class=\"wp-block-heading\">Expose metadata to training pipelines via a manifest<\/h3>\n\n\n\n<p class=\"wp-block-paragraph\">Rather than having training jobs scan bucket prefixes directly, generate a manifest file (JSONL or Parquet) at the end of each ingest batch that enumerates every object key + its full metadata. Tools like PyTorch&#8217;s WebDataset and HuggingFace datasets can load directly from these manifests, and it decouples the training job from needing object storage credentials for discovery.<\/p>\n\n\n\n<h2 class=\"wp-block-heading\">Deciding on storage while evaluating data<\/h2>\n\n\n\n<p class=\"wp-block-paragraph\">To summarize, this is why choosing your storage destination based on your current (or, if you\u2019re already undergoing a data management transformation, future-state) scenario is critical:<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li><strong>Ingest type: <\/strong>Does your storage provide enough network bandwidth to capture the correct type of data in real time (especially for video?)<\/li>\n\n\n\n<li><strong>Data type: <\/strong>Will large files potentially incur large upfront costs with upload fees?<\/li>\n\n\n\n<li><strong>Data processing workflow: <\/strong>Does your storage have capacity headroom for file multiplication during processing, and does the cost structure allow this to happen without potentially draining infrastructure budgets?<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Even with Backblaze B2\u2019s hot storage at cold storage pricing, it may be beneficial for you to tier your storage based on your ingestion type, and how frequently the data will be accessed for training.<\/p>\n\n\n\n<ul class=\"wp-block-list\">\n<li>Batch ingestion is better suited for mid to lower performance storage, as this is typically used for historical datasets or a set schedule of pre-determined data updates, such as jobs pulling from relational databases or CSV uploads once a day or once per week.<\/li>\n\n\n\n<li>Streaming ingestion is well-suited for hot storage to support a continuous stream of real-time (or near-real-time) data processing, such as from social media feeds and high-volume e-commerce AI helper agents.<\/li>\n\n\n\n<li>Hybrid ingestion uses a combination of batch and streaming ingestion to handle both historical and real-time data requirements for AI models.<\/li>\n<\/ul>\n\n\n\n<p class=\"wp-block-paragraph\">Surprise retrieval penalties for model training can happen even while building out your MVP or proof of concept\u2013so avoid a storage headache and start building your pipeline at data ingest with Backblaze B2 $6.95 per TB\/month. <a href=\"https:\/\/www.backblaze.com\/sign-up\/ai-cloud-storage\" target=\"_blank\" rel=\"noreferrer noopener\">Create an account<\/a> to get started with 10GB for free, or <a href=\"https:\/\/www.backblaze.com\/contact-sales\/cloud-storage\/\" target=\"_blank\" rel=\"noreferrer noopener\">contact our storage experts<\/a> for assistance with migrations and more.<\/p>\n\n\n\n<p class=\"wp-block-paragraph\"><strong>Looking for more info on data ingest?<\/strong> Watch the <a href=\"https:\/\/www.brighttalk.com\/webcast\/14807\/672378?utm_source=brighttalk-portal&amp;utm_medium=web&amp;utm_campaign=channel-page&amp;utm_content=recorded\" target=\"_blank\" rel=\"noreferrer noopener\">on-demand webinar<\/a> that dives into more details about data ingest with Backblaze\u2019s Director of Applied AI Jeronimo De Leon.<\/p>\n\n\n<div class=\"wp-block-image\">\n<figure class=\"aligncenter size-full\"><a href=\"https:\/\/www.brighttalk.com\/webcast\/14807\/672378?utm_source=brighttalk-portal&amp;utm_medium=web&amp;utm_campaign=channel-page&amp;utm_content=recorded\" target=\"_blank\" rel=\" noreferrer noopener\"><img loading=\"lazy\" decoding=\"async\" width=\"640\" height=\"360\" src=\"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/image_1101897.webp\" alt=\"A webinar title card.\" class=\"wp-image-113282\" srcset=\"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/image_1101897.webp 640w, https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/image_1101897-300x169.webp 300w\" sizes=\"auto, (max-width: 640px) 100vw, 640px\" \/><\/a><\/figure>\n<\/div>\n\n\n<div style=\"height:15px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<p class=\"wp-block-paragraph\"><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Most AI projects don&#8217;t fail because of the model\u2014they fail because of the data. Learn how to build an AI-ready data pipeline from ingestion onward.<\/p>\n","protected":false},"author":224,"featured_media":113277,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"content-type":"","footnotes":"","jetpack_post_was_ever_published":false},"categories":[7,434,438],"tags":[489,468],"class_list":["post-113274","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-cloud-storage","category-featured-1","category-featured-cloud-storage","tag-ai-ml","tag-b2cloud","entry"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.1 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>AI Data Pipeline 101: Ingest &amp; Archive<\/title>\n<meta name=\"description\" content=\"Build better AI from the start. Learn best practices for data pipelines that improve model performance and reliability.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"AI Data Pipeline 101: Ingest &amp; Archive\" \/>\n<meta property=\"og:description\" content=\"Build better AI from the start. Learn best practices for data pipelines that improve model performance and reliability.\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/\" \/>\n<meta property=\"og:site_name\" content=\"Backblaze Blog | Cloud Storage &amp; Cloud Backup\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/backblaze\" \/>\n<meta property=\"article:published_time\" content=\"2026-07-29T14:25:44+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-07-29T14:25:46+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/AI-0002-Blog-Header-1440x820-1.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1440\" \/>\n\t<meta property=\"og:image:height\" content=\"820\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Maddie Presland\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@backblaze\" \/>\n<meta name=\"twitter:site\" content=\"@backblaze\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Maddie Presland\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"7 minutes\" \/>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"AI Data Pipeline 101: Ingest &amp; Archive","description":"Build better AI from the start. Learn best practices for data pipelines that improve model performance and reliability.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/","og_locale":"en_US","og_type":"article","og_title":"AI Data Pipeline 101: Ingest &amp; Archive","og_description":"Build better AI from the start. Learn best practices for data pipelines that improve model performance and reliability.","og_url":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/","og_site_name":"Backblaze Blog | Cloud Storage &amp; Cloud Backup","article_publisher":"https:\/\/www.facebook.com\/backblaze","article_published_time":"2026-07-29T14:25:44+00:00","article_modified_time":"2026-07-29T14:25:46+00:00","og_image":[{"width":1440,"height":820,"url":"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/AI-0002-Blog-Header-1440x820-1.png","type":"image\/png"}],"author":"Maddie Presland","twitter_card":"summary_large_image","twitter_creator":"@backblaze","twitter_site":"@backblaze","twitter_misc":{"Written by":"Maddie Presland","Est. reading time":"7 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/#article","isPartOf":{"@id":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/"},"author":{"name":"Maddie Presland","@id":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/#\/schema\/person\/5a95887c8e781ea9cf10472e47175ce0"},"headline":"AI Data Pipeline 101: Ingest &amp; Archive","datePublished":"2026-07-29T14:25:44+00:00","dateModified":"2026-07-29T14:25:46+00:00","mainEntityOfPage":{"@id":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/"},"wordCount":1432,"commentCount":0,"publisher":{"@id":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/#organization"},"image":{"@id":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/#primaryimage"},"thumbnailUrl":"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/AI-0002-Blog-Header-1440x820-1.png","keywords":["AI\/ML","B2Cloud"],"articleSection":["Cloud Storage","Featured","Featured-Cloud Storage"],"inLanguage":"en-US","potentialAction":[{"@type":"CommentAction","name":"Comment","target":["https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/#respond"]}]},{"@type":"WebPage","@id":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/","url":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/","name":"AI Data Pipeline 101: Ingest &amp; Archive","isPartOf":{"@id":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/#primaryimage"},"image":{"@id":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/#primaryimage"},"thumbnailUrl":"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/AI-0002-Blog-Header-1440x820-1.png","datePublished":"2026-07-29T14:25:44+00:00","dateModified":"2026-07-29T14:25:46+00:00","description":"Build better AI from the start. Learn best practices for data pipelines that improve model performance and reliability.","breadcrumb":{"@id":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/#primaryimage","url":"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/AI-0002-Blog-Header-1440x820-1.png","contentUrl":"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/AI-0002-Blog-Header-1440x820-1.png","width":1440,"height":820,"caption":"An illustration of gears, boxes, and a graphic that says AI."},{"@type":"BreadcrumbList","@id":"https:\/\/www.backblaze.com\/blog\/ai-data-pipeline-101-ingest-archive\/#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/"},{"@type":"ListItem","position":2,"name":"AI Data Pipeline 101: Ingest &amp; Archive"}]},{"@type":"WebSite","@id":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/#website","url":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/","name":"Backblaze Cloud Solutions Blog","description":"Cloud Storage &amp; Cloud Backup","publisher":{"@id":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/#organization","name":"Backblaze","url":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/#\/schema\/logo\/image\/","url":"https:\/\/i0.wp.com\/www.backblaze.com\/blog\/wp-content\/uploads\/2017\/12\/backblaze_icon_transparent.png?fit=512%2C512&ssl=1","contentUrl":"https:\/\/i0.wp.com\/www.backblaze.com\/blog\/wp-content\/uploads\/2017\/12\/backblaze_icon_transparent.png?fit=512%2C512&ssl=1","width":512,"height":512,"caption":"Backblaze"},"image":{"@id":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/backblaze","https:\/\/x.com\/backblaze","https:\/\/www.youtube.com\/user\/Backblaze","https:\/\/en.wikipedia.org\/wiki\/Backblaze"]},{"@type":"Person","@id":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/#\/schema\/person\/5a95887c8e781ea9cf10472e47175ce0","name":"Maddie Presland","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2025\/10\/Backblaze_Author-Maddie-Presland_Square-150x150.jpg","url":"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2025\/10\/Backblaze_Author-Maddie-Presland_Square-150x150.jpg","contentUrl":"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2025\/10\/Backblaze_Author-Maddie-Presland_Square-150x150.jpg","caption":"Maddie Presland"},"description":"Maddie Presland is a Product Marketing Manager at Backblaze specializing in app storage use cases for multi-cloud architectures and AI. Maddie has more than five years of experience as a product marketer focusing on cloud infrastructure and developing technical marketing content for developers. With a background in journalism, she combines storytelling with her technical curiosity and ability to crash course just about anything. Connect with her on LinkedIn.","url":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/author\/maddiepresland\/"}]}},"jetpack_featured_media_url":"https:\/\/backblazeprod.wpenginepowered.com\/wp-content\/uploads\/2026\/07\/AI-0002-Blog-Header-1440x820-1.png","_links":{"self":[{"href":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/wp-json\/wp\/v2\/posts\/113274","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/wp-json\/wp\/v2\/users\/224"}],"replies":[{"embeddable":true,"href":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/wp-json\/wp\/v2\/comments?post=113274"}],"version-history":[{"count":5,"href":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/wp-json\/wp\/v2\/posts\/113274\/revisions"}],"predecessor-version":[{"id":113285,"href":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/wp-json\/wp\/v2\/posts\/113274\/revisions\/113285"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/wp-json\/wp\/v2\/media\/113277"}],"wp:attachment":[{"href":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/wp-json\/wp\/v2\/media?parent=113274"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/wp-json\/wp\/v2\/categories?post=113274"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/backblazeprod.wpenginepowered.com\/blog\/wp-json\/wp\/v2\/tags?post=113274"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}