{"id":17637,"date":"2015-02-28T18:05:28","date_gmt":"2015-02-28T12:35:28","guid":{"rendered":"https:\/\/2thenew.online\/blog\/?p=17637"},"modified":"2024-01-02T17:49:28","modified_gmt":"2024-01-02T12:19:28","slug":"apache-flume-setup-best-practices","status":"publish","type":"post","link":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/","title":{"rendered":"Apache Flume : Setup &amp; Best Practices"},"content":{"rendered":"<p>Apache Flume is an open source project aimed at providing a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large volume of data. It is a complex task when moving data in large volume. We try to minimize the latency in transfer; this is achieved by specifically tweaking the configuration of Flume. First, we&#8217;ll see how to setup flume.<\/p>\n<p><em><strong>Setting Up Flume :<\/strong><\/em><\/p>\n<ol>\n<li>Download flume binary from <a title=\"Flume Download\" href=\"http:\/\/flume.apache.org\/download.html\">http:\/\/flume.apache.org\/download.html<\/a><\/li>\n<li>Extract and put the binary folder in globally accessible place. For e.g. \/usr\/local\/flume\u00a0\u00a0 (Use &#8220;<em>tar -xvf apache-flume-1.5.2-bin.tar.gz&#8221; <\/em>and <em>&#8220;mv apache-flume-1.5.2-bin \/usr\/local\/flume&#8221;<\/em>)<\/li>\n<li>Once done, set the global variables in the &#8220;<em>.bashrc<\/em>&#8221; file of the user accessing Flume.<a href=\"\/blog\/wp-ttn-blog\/uploads\/2015\/02\/Screenshot-from-2015-02-27-151944.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter  wp-image-17640\" src=\"\/blog\/wp-ttn-blog\/uploads\/2015\/02\/Screenshot-from-2015-02-27-151944.png\" alt=\"Flume Variables\" width=\"401\" height=\"134\" \/><\/a><\/li>\n<li>Use &#8220;<em>source .bashrc&#8221; <\/em>to set the new variables in effect. Test by running command &#8220;<em>flume-ng version&#8221;.<\/em><\/li>\n<li>Once the command shows the right version, Flume is set to work on the current system. Please note &#8220;<em>flume-ng&#8221;<\/em> represents, &#8220;<strong>Flume Next-Gen<\/strong>&#8220;<\/li>\n<\/ol>\n<p>&nbsp;<\/p>\n<p>Flume Agent is the one which takes care of the whole process of taking data from &#8220;<strong><em>source<\/em><\/strong>&#8220;, putting it on to the &#8220;<em><strong>channel<\/strong><\/em>&#8220;, and finally dumping it in the &#8220;<em><strong>sink<\/strong><\/em>&#8220;. &#8220;<em><strong>Sink<\/strong><\/em>&#8221; usually is <em>HDFS.<\/em><\/p>\n<img loading=\"lazy\" decoding=\"async\" class=\"aligncenter\" src=\"\/blog\/wp-ttn-blog\/uploads\/2024\/01\/DevGuide_image00.png\" alt=\"Blog image\" width=\"520\" height=\"218\" \/>\n<p><em><strong>Basic Configuration<\/strong><\/em> :<\/p>\n<blockquote><p><em>Apache Flume takes a configuration file every time it runs a task. This task is kept alive in order to listen to any change in the <strong>source<\/strong>, and must be terminated manually by the user. A basic configuration to read data taking &#8220;<strong>local file system&#8221;<\/strong> as the source, keeping channel as &#8220;<\/em><strong>memory<\/strong>&#8220;, and &#8220;<em><strong>hdfs<\/strong><\/em>&#8221; as sink, could be:<\/p>\n<p>&nbsp;<\/p><\/blockquote>\n<p><a href=\"\/blog\/wp-ttn-blog\/uploads\/2015\/02\/Screenshot-from-2015-02-27-1519441.png\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-full wp-image-17648\" src=\"\/blog\/wp-ttn-blog\/uploads\/2015\/02\/Screenshot-from-2015-02-27-1519441.png\" alt=\"Flume Config\" width=\"730\" height=\"476\" \/><\/a><\/p>\n<p>Sources could be anything like an Avro Client being used by Log4JAppender as well.<\/p>\n<p>It is worth noting, the above given configuration will produce separate files of 10 records each by default, taking timestamp ( <em><strong>an interceptor <\/strong><\/em>) as the point of reference of last update of the source.<\/p>\n<p>The file so created be saved as flume.conf, and can be run as :<\/p>\n<p><em><strong>flume-ng agent -f flume.conf -n source_agent<\/strong><\/em><\/p>\n<p><em><strong>Performance Measures, Issues and Comments:<\/strong><\/em><\/p>\n<ol>\n<li>Memory leaks in <strong>log4jappender<\/strong> when ingesting just 10000 records.\n<ul>\n<li>Solved by making thread sleep after every 500 records, thus decreasing the load on the channel<\/li>\n<\/ul>\n<\/li>\n<li>\u00a0GC Memory Leak (Flume level)\n<ul>\n<li>Solved by keeping transactionCapacity of channel low and capacity of channel high enough<\/li>\n<\/ul>\n<\/li>\n<li>Avro&#8217;s Optimized data serialization is not expolited when using log4jappender\n<ul>\n<li>Solved by directly pushing the file onto avro-client (inbuilt in flume-ng)<\/li>\n<\/ul>\n<\/li>\n<li>Due to memory leaks, could not write more than 12000 records from log4jappender\n<ul>\n<li>Solved in points 1 &amp; 2<\/li>\n<\/ul>\n<\/li>\n<li>Did performance analysis of data ingestion directly from file as well as avro client of 14000 records when only 10 records rolled up per file in HDFS\n<ul>\n<li>Using Avro Client\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0\u00a0 \u00a0\u00a0 &#8211;&gt; Total time &#8211; <strong>00:02:48<\/strong><\/li>\n<li>Directly from File System &#8211;&gt; Total time &#8211; <strong>00:02:44<\/strong><\/li>\n<\/ul>\n<\/li>\n<li>When using avro client, realtime update to the file was not taken into account\n<ul>\n<li>Unsolved problem<\/li>\n<\/ul>\n<\/li>\n<li>When ingesting directly from file, realtime updates were automatically registered, on the basis of timestamp for last modification\n<ul>\n<li>Used the following configuration<br \/>\nsource_agent.sources.test_source.interceptors = itime<br \/>\n# http:\/\/flume.apache.org\/FlumeUserGuide.html#timestamp-interceptor<br \/>\nsource_agent.sources.test_source.interceptors.itime.type = timestamp<\/li>\n<\/ul>\n<\/li>\n<li>Flume could only create files to write at max 10 records in a single file by default. This decreased the ingestion rate, thus increasing the ingestion time\n<ul>\n<li>Solved by changing the rollCount property of the HDFSSink to the desired number of events\/records per file; this increased the ingestion rate, thus decreasing the ingestion time.<\/li>\n<li>Also made rollSize and rollInterval as 0, so that they are not used.<\/li>\n<\/ul>\n<\/li>\n<li>While controlling the channel capacity, encountered Memory Leaks\n<ul>\n<li>Solved by changing the rollCount(HDFSSink) and transactionCapacity(channel) so that the channel is cleared for more data<\/li>\n<\/ul>\n<\/li>\n<li>Did performance analysis of data ingestion directly from file after applying solutions from 8 &amp; 9 of 20000 records\n<ul>\n<li>Total time &#8211; <strong>00:00:04 (<\/strong>Previously around <strong>00:02:30)<\/strong><\/li>\n<\/ul>\n<\/li>\n<li>How to trigger notification on HDFS update from flume\n<ul>\n<li>Can be potentially solved by Oozie Coordinator<\/li>\n<\/ul>\n<\/li>\n<\/ol>\n<p>&nbsp;<\/p>\n<p>This covers the very basics of setting up Flume and mitigating some of the common issues which one encounters while using it.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Apache Flume is an open source project aimed at providing a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large volume of data. It is a complex task when moving data in large volume. We try to minimize the latency in transfer; this is achieved by specifically tweaking the configuration of Flume. [&hellip;]<\/p>\n","protected":false},"author":158,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"iawp_total_views":14,"footnotes":""},"categories":[1395],"tags":[],"class_list":["post-17637","post","type-post","status-publish","format-standard","hentry","category-big-data"],"aioseo_notices":[],"aioseo_head":"\n\t\t<!-- All in One SEO 5.0.0.1 - aioseo.com -->\n\t<meta name=\"description\" content=\"Apache Flume is an open source project aimed at providing a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large volume of data. It is a complex task when moving data in large volume. We try to minimize the latency in transfer; this is achieved by specifically tweaking the configuration of Flume.\" \/>\n\t<meta name=\"robots\" content=\"max-image-preview:large\" \/>\n\t<meta name=\"author\" content=\"Rishabh Jain\"\/>\n\t<link rel=\"canonical\" href=\"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/\" \/>\n\t<meta name=\"generator\" content=\"All in One SEO (AIOSEO) 5.0.0.1\" \/>\n\t\t<meta property=\"og:locale\" content=\"en_US\" \/>\n\t\t<meta property=\"og:site_name\" content=\"TO THE NEW BLOG\" \/>\n\t\t<meta property=\"og:type\" content=\"blog\" \/>\n\t\t<meta property=\"og:title\" content=\"Apache Flume : Setup &amp; Best Practices | TO THE NEW Blog\" \/>\n\t\t<meta property=\"og:description\" content=\"Apache Flume is an open source project aimed at providing a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large volume of data. It is a complex task when moving data in large volume. We try to minimize the latency in transfer; this is achieved by specifically tweaking the configuration of Flume.\" \/>\n\t\t<meta property=\"og:url\" content=\"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/\" \/>\n\t\t<meta property=\"og:image\" content=\"https:\/\/2thenew.online\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png\" \/>\n\t\t<meta property=\"og:image:secure_url\" content=\"https:\/\/2thenew.online\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png\" \/>\n\t\t<meta name=\"twitter:card\" content=\"summary\" \/>\n\t\t<meta name=\"twitter:site\" content=\"@tothenew\" \/>\n\t\t<meta name=\"twitter:title\" content=\"Apache Flume : Setup &amp; Best Practices | TO THE NEW Blog\" \/>\n\t\t<meta name=\"twitter:description\" content=\"Apache Flume is an open source project aimed at providing a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large volume of data. It is a complex task when moving data in large volume. We try to minimize the latency in transfer; this is achieved by specifically tweaking the configuration of Flume.\" \/>\n\t\t<meta name=\"twitter:image\" content=\"https:\/\/2thenew.online\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png\" \/>\n\t\t<script type=\"application\/ld+json\" class=\"aioseo-schema\">\n\t\t\t{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/apache-flume-setup-best-practices\\\/#article\",\"name\":\"Apache Flume : Setup & Best Practices | TO THE NEW Blog\",\"headline\":\"Apache Flume : Setup &amp; Best Practices\",\"author\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/rishabhj\\\/#author\"},\"publisher\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#organization\"},\"image\":{\"@type\":\"ImageObject\",\"url\":\"\\\/blog\\\/wp-ttn-blog\\\/uploads\\\/2015\\\/02\\\/Screenshot-from-2015-02-27-151944.png\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/apache-flume-setup-best-practices\\\/#articleImage\"},\"datePublished\":\"2015-02-28T18:05:28+05:30\",\"dateModified\":\"2024-01-02T17:49:28+05:30\",\"inLanguage\":\"en-US\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/apache-flume-setup-best-practices\\\/#webpage\"},\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/apache-flume-setup-best-practices\\\/#webpage\"},\"articleSection\":\"Big Data\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/apache-flume-setup-best-practices\\\/#breadcrumblist\",\"itemListElement\":[{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog#listItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.tothenew.com\\\/blog\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/category\\\/big-data\\\/#listItem\",\"name\":\"Big Data\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/category\\\/big-data\\\/#listItem\",\"position\":2,\"name\":\"Big Data\",\"item\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/category\\\/big-data\\\/\",\"nextItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/apache-flume-setup-best-practices\\\/#listItem\",\"name\":\"Apache Flume : Setup &amp; Best Practices\"},\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog#listItem\",\"name\":\"Home\"}},{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/apache-flume-setup-best-practices\\\/#listItem\",\"position\":3,\"name\":\"Apache Flume : Setup &amp; Best Practices\",\"previousItem\":{\"@type\":\"ListItem\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/category\\\/big-data\\\/#listItem\",\"name\":\"Big Data\"}}]},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#organization\",\"name\":\"TO THE NEW Blog\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/\"},{\"@type\":\"Person\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/rishabhj\\\/#author\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/rishabhj\\\/\",\"name\":\"Rishabh Jain\",\"image\":{\"@type\":\"ImageObject\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/apache-flume-setup-best-practices\\\/#authorImage\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/3104b9b638339a2e4703ce6db185dcc56367d6ee39d2cab8c3069eda20f7350b?s=96&d=mm&r=g\",\"width\":96,\"height\":96,\"caption\":\"Rishabh Jain\"}},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/apache-flume-setup-best-practices\\\/#webpage\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/apache-flume-setup-best-practices\\\/\",\"name\":\"Apache Flume : Setup & Best Practices | TO THE NEW Blog\",\"description\":\"Apache Flume is an open source project aimed at providing a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large volume of data. It is a complex task when moving data in large volume. We try to minimize the latency in transfer; this is achieved by specifically tweaking the configuration of Flume.\",\"inLanguage\":\"en-US\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#website\"},\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/apache-flume-setup-best-practices\\\/#breadcrumblist\"},\"author\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/rishabhj\\\/#author\"},\"creator\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/author\\\/rishabhj\\\/#author\"},\"datePublished\":\"2015-02-28T18:05:28+05:30\",\"dateModified\":\"2024-01-02T17:49:28+05:30\"},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#website\",\"url\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/\",\"name\":\"TO THE NEW Blog\",\"inLanguage\":\"en-US\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.tothenew.com\\\/blog\\\/#organization\"}}]}\n\t\t<\/script>\n\t\t<!-- All in One SEO -->\n\n","aioseo_head_json":{"title":"Apache Flume : Setup & Best Practices | TO THE NEW Blog","description":"Apache Flume is an open source project aimed at providing a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large volume of data. It is a complex task when moving data in large volume. We try to minimize the latency in transfer; this is achieved by specifically tweaking the configuration of Flume.","canonical_url":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/","robots":"max-image-preview:large","keywords":"","webmasterTools":{"miscellaneous":""},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/#article","name":"Apache Flume : Setup & Best Practices | TO THE NEW Blog","headline":"Apache Flume : Setup &amp; Best Practices","author":{"@id":"https:\/\/2thenew.online\/blog\/author\/rishabhj\/#author"},"publisher":{"@id":"https:\/\/2thenew.online\/blog\/#organization"},"image":{"@type":"ImageObject","url":"\/blog\/wp-ttn-blog\/uploads\/2015\/02\/Screenshot-from-2015-02-27-151944.png","@id":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/#articleImage"},"datePublished":"2015-02-28T18:05:28+05:30","dateModified":"2024-01-02T17:49:28+05:30","inLanguage":"en-US","mainEntityOfPage":{"@id":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/#webpage"},"isPartOf":{"@id":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/#webpage"},"articleSection":"Big Data"},{"@type":"BreadcrumbList","@id":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/#breadcrumblist","itemListElement":[{"@type":"ListItem","@id":"https:\/\/2thenew.online\/blog#listItem","position":1,"name":"Home","item":"https:\/\/2thenew.online\/blog","nextItem":{"@type":"ListItem","@id":"https:\/\/2thenew.online\/blog\/category\/big-data\/#listItem","name":"Big Data"}},{"@type":"ListItem","@id":"https:\/\/2thenew.online\/blog\/category\/big-data\/#listItem","position":2,"name":"Big Data","item":"https:\/\/2thenew.online\/blog\/category\/big-data\/","nextItem":{"@type":"ListItem","@id":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/#listItem","name":"Apache Flume : Setup &amp; Best Practices"},"previousItem":{"@type":"ListItem","@id":"https:\/\/2thenew.online\/blog#listItem","name":"Home"}},{"@type":"ListItem","@id":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/#listItem","position":3,"name":"Apache Flume : Setup &amp; Best Practices","previousItem":{"@type":"ListItem","@id":"https:\/\/2thenew.online\/blog\/category\/big-data\/#listItem","name":"Big Data"}}]},{"@type":"Organization","@id":"https:\/\/2thenew.online\/blog\/#organization","name":"TO THE NEW Blog","url":"https:\/\/2thenew.online\/blog\/"},{"@type":"Person","@id":"https:\/\/2thenew.online\/blog\/author\/rishabhj\/#author","url":"https:\/\/2thenew.online\/blog\/author\/rishabhj\/","name":"Rishabh Jain","image":{"@type":"ImageObject","@id":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/#authorImage","url":"https:\/\/secure.gravatar.com\/avatar\/3104b9b638339a2e4703ce6db185dcc56367d6ee39d2cab8c3069eda20f7350b?s=96&d=mm&r=g","width":96,"height":96,"caption":"Rishabh Jain"}},{"@type":"WebPage","@id":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/#webpage","url":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/","name":"Apache Flume : Setup & Best Practices | TO THE NEW Blog","description":"Apache Flume is an open source project aimed at providing a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large volume of data. It is a complex task when moving data in large volume. We try to minimize the latency in transfer; this is achieved by specifically tweaking the configuration of Flume.","inLanguage":"en-US","isPartOf":{"@id":"https:\/\/2thenew.online\/blog\/#website"},"breadcrumb":{"@id":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/#breadcrumblist"},"author":{"@id":"https:\/\/2thenew.online\/blog\/author\/rishabhj\/#author"},"creator":{"@id":"https:\/\/2thenew.online\/blog\/author\/rishabhj\/#author"},"datePublished":"2015-02-28T18:05:28+05:30","dateModified":"2024-01-02T17:49:28+05:30"},{"@type":"WebSite","@id":"https:\/\/2thenew.online\/blog\/#website","url":"https:\/\/2thenew.online\/blog\/","name":"TO THE NEW Blog","inLanguage":"en-US","publisher":{"@id":"https:\/\/2thenew.online\/blog\/#organization"}}]},"og:locale":"en_US","og:site_name":"TO THE NEW BLOG","og:type":"blog","og:title":"Apache Flume : Setup &amp; Best Practices | TO THE NEW Blog","og:description":"Apache Flume is an open source project aimed at providing a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large volume of data. It is a complex task when moving data in large volume. We try to minimize the latency in transfer; this is achieved by specifically tweaking the configuration of Flume.","og:url":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/","og:image":"https:\/\/2thenew.online\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png","og:image:secure_url":"https:\/\/2thenew.online\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png","twitter:card":"summary","twitter:site":"@tothenew","twitter:title":"Apache Flume : Setup &amp; Best Practices | TO THE NEW Blog","twitter:description":"Apache Flume is an open source project aimed at providing a distributed, reliable, and available service for efficiently collecting, aggregating, and moving large volume of data. It is a complex task when moving data in large volume. We try to minimize the latency in transfer; this is achieved by specifically tweaking the configuration of Flume.","twitter:image":"https:\/\/2thenew.online\/blog\/wp-content\/themes\/ttn\/images\/social-logo.png"},"aioseo_meta_data":{"post_id":"17637","title":"Apache Flume : Setup &amp; Best Practices | #site_title","description":null,"keywords":null,"keyphrases":null,"primary_term":null,"canonical_url":null,"og_title":null,"og_description":null,"og_object_type":"default","og_image_type":"default","og_image_url":null,"og_image_width":null,"og_image_height":null,"og_image_custom_url":null,"og_image_custom_fields":null,"og_video":null,"og_custom_url":null,"og_article_section":null,"og_article_tags":null,"twitter_use_og":false,"twitter_card":"default","twitter_image_type":"default","twitter_image_url":null,"twitter_image_custom_url":null,"twitter_image_custom_fields":null,"twitter_title":null,"twitter_description":null,"schema":{"blockGraphs":[],"customGraphs":[],"default":{"data":{"Article":[],"Course":[],"Dataset":[],"FAQPage":[],"Movie":[],"Person":[],"Product":[],"ProductReview":[],"Car":[],"Recipe":[],"Service":[],"SoftwareApplication":[],"WebPage":[]},"graphName":"Article","isEnabled":true},"graphs":[]},"schema_type":null,"schema_type_options":null,"pillar_content":false,"robots_default":true,"robots_noindex":false,"robots_noarchive":false,"robots_nosnippet":false,"robots_nofollow":false,"robots_noimageindex":false,"robots_noodp":false,"robots_notranslate":false,"robots_max_snippet":null,"robots_max_videopreview":null,"robots_max_imagepreview":"large","priority":null,"frequency":null,"local_seo":null,"limit_modified_date":false,"created":"2021-04-30 07:44:45","updated":"2024-02-29 10:39:22","focus_keyword":null,"additional_keywords":null,"truseo_locale":null,"ai":null,"breadcrumb_settings":null,"seo_analyzer_scan_date":null},"aioseo_breadcrumb":"<div class=\"aioseo-breadcrumbs\"><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/2thenew.online\/blog\" title=\"Home\">Home<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\t<a href=\"https:\/\/2thenew.online\/blog\/category\/big-data\/\" title=\"Big Data\">Big Data<\/a>\n\t\t<\/span><span class=\"aioseo-breadcrumb-separator\">&raquo;<\/span><span class=\"aioseo-breadcrumb\">\n\t\t\tApache Flume : Setup &amp; Best Practices\n\t\t<\/span><\/div>","aioseo_breadcrumb_json":[{"label":"Home","link":"https:\/\/2thenew.online\/blog"},{"label":"Big Data","link":"https:\/\/2thenew.online\/blog\/category\/big-data\/"},{"label":"Apache Flume : Setup &amp; Best Practices","link":"https:\/\/2thenew.online\/blog\/apache-flume-setup-best-practices\/"}],"_links":{"self":[{"href":"https:\/\/2thenew.online\/blog\/wp-json\/wp\/v2\/posts\/17637","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/2thenew.online\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/2thenew.online\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/2thenew.online\/blog\/wp-json\/wp\/v2\/users\/158"}],"replies":[{"embeddable":true,"href":"https:\/\/2thenew.online\/blog\/wp-json\/wp\/v2\/comments?post=17637"}],"version-history":[{"count":1,"href":"https:\/\/2thenew.online\/blog\/wp-json\/wp\/v2\/posts\/17637\/revisions"}],"predecessor-version":[{"id":59899,"href":"https:\/\/2thenew.online\/blog\/wp-json\/wp\/v2\/posts\/17637\/revisions\/59899"}],"wp:attachment":[{"href":"https:\/\/2thenew.online\/blog\/wp-json\/wp\/v2\/media?parent=17637"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/2thenew.online\/blog\/wp-json\/wp\/v2\/categories?post=17637"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/2thenew.online\/blog\/wp-json\/wp\/v2\/tags?post=17637"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}