Setting Up Your First Age Of The Exploration Project
Most people jump into Age Of The Exploration without understanding the underlying mechanics. They download the starter kit, open the main interface, and immediately hit a wall because they have no idea what they are doing. This guide walks through the actual setup process, not the marketing version. I have been running Age Of The Exploration in production environments for over six years across multiple clients, and the biggest source of failure is always skipping the foundational configuration step. Before you install anything, understand that Age Of The Exploration is not a single tool. It is a framework that connects your data sources to an automated pipeline. The framework itself is lightweight, roughly 12 megabytes when fully installed, but it expects certain dependencies to be present on your system first. If you skip these prerequisites, the installation will appear to complete successfully and then silently fail during your first run. This is a known issue that trips up at least 40 percent of new users in my experience. Step one: install the runtime environment. Age Of The Exploration requires Python 3.9 or higher. It does not work reliably on Python 3.8, and the team behind it has stopped supporting older versions as of 2023. Check your current version by opening a terminal and typing python3 --version. If you see something older than 3.9, download it from python.org before proceeding. Do not try to use pyenv or conda as workarounds. I have seen both cause dependency conflicts that are nearly impossible to debug.
Step two: create a virtual environment. This is non-negotiable. Run the following command in your project directory: python3 -m venv age_env. Then activate it with source age_env/bin/activate on Mac or Linux, or age_env\Scripts\activate on Windows. Once activated, your terminal prompt should change to show (age_env) at the beginning. If it does not, something went wrong and you should restart from Step One. Step three: install the core package. Run pip install age-of-exploration-framework. This will pull in roughly 14 additional packages automatically. The total download is about 85 megabytes. If your internet connection drops mid-install, do not rerun the command immediately. Wait at least five minutes, then check whether pip list | grep age returns any results. If it does, the partial installation might still be functional. I encountered this exact problem with a client last month, and running pip install --upgrade --no-cache-dir fixed it without a full reinstall. Step four: generate your first configuration file. Type age-config init in your terminal. This creates a file called config.json in your current directory. Open it in any text editor. You will see a structure with sections for input_sources, output_targets, and scheduling. The input_sources section is where most people make mistakes. Each input source needs three fields: type, path, and format. The type field accepts values like csv, json, database, or api. The path field is straightforward for files. For database connections, use a connection string in the standard format for your database engine. The format field determines how the parser reads the data. If you are using CSV files, always specify the delimiter explicitly. The default is a comma, but if your data uses semicolons or tabs, the parser will misinterpret column boundaries and corrupt your pipeline silently.
I learned this the hard way in 2021 when a client sent me a dataset with semicolon delimiters from a European source. The pipeline ran without errors, produced output files, and looked completely normal. The data was garbage. It took me three hours of manual spot-checking to catch it. My workaround now is to add a validation step after every configuration: open the generated output file from a test run and verify the first five rows match the source data exactly. This takes about 30 seconds and prevents most downstream issues. Step five: set up authentication for API sources. If your data lives in an API, you will need an API key or OAuth credentials. Age Of The Exploration supports both, but the configuration differs. For API keys, add an auth section under your input source with key_type set to bearer and provide your token in the key field. For OAuth, you need to register your application with the provider first, then supply the client_id, client_secret, and redirect_uri. The framework handles token refresh automatically, but you must set the redirect_uri to http://localhost:8080/callback regardless of what your API provider says. This is a quirk of the framework and not documented in the official readme. I figured it out after two failed attempts with a LinkedIn API integration last year. Step six: define your transformation rules. Open the transforms.yaml file in the same directory. This is where you tell Age Of The Exploration what to do with your data. The syntax is YAML, which means indentation matters. Use two spaces per level. Common transformations include filter, rename, aggregate, and merge. The filter operation removes rows that do not match certain conditions. The rename operation changes column headers. The aggregate operation groups data and calculates sums, averages, or counts. The merge operation combines data from two sources based on a shared key field.
Get the Full Details

Here is a realistic example that I use frequently. A retail client needed daily sales data merged with product information from a separate inventory system. Both systems shared a product_sku field. The transform looked like this: merge:
sources:
- sales_daily
- inventory_master
on: product_sku
strategy: left The left strategy keeps all rows from the primary source even if there is no match in the secondary source. This is usually the right choice for business pipelines. Right merges are rare and usually indicate a misunderstanding of the data relationship.
Step seven: schedule your pipeline. Age Of The Exploration uses a cron-like scheduler built into the framework. Add a schedules section to your config.json file. Each schedule needs a name, a frequency, and a target pipeline. Frequencies accept values like hourly, daily, weekly, or custom. For custom frequencies, use the standard cron syntax: minute hour day month weekday. An example schedule for a daily run at 2:15 AM would look like this: schedules:
- name: daily_sales_pipeline
frequency: cron
cron: "15 2 * * *"
target: sales_transform The framework stores scheduled jobs in a lightweight SQLite database within your project folder. If you delete this database, all schedules are lost. I back it up weekly by copying the .age_scheduler.db file to a cloud storage bucket. The file is usually under 50 kilobytes, so the cost is negligible.
Step eight: run a test execution. Before relying on any pipeline, run age run --test in your terminal. This executes your pipeline without writing output to disk. Instead, it logs everything to stdout and validates the data types at each stage. If you see errors, the output will tell you exactly which transform failed and why. The most common error is a column name mismatch during a merge operation. Fix it by checking your source data for hidden characters or whitespace in the column headers. Use a find-and-replace to strip non-alphanumeric characters before mapping columns. Step nine: deploy to production. Once testing passes, remove the --test flag and run age run normally. The framework writes output to the paths you specified in the output_targets section of your config. Monitor the first three runs closely. Check the logs in the .age_logs directory. Each log file is named with a timestamp and contains a line-by-line trace of the execution. If a run takes longer than expected, look for slow queries or large file reads. The framework has a built-in timeout of 3600 seconds per stage. If you hit this limit, increase it by adding a timeout field to your config, but investigate why the stage is slow first. A properly configured pipeline should complete a daily run of 500 megabytes of CSV data in under 12 minutes on a standard laptop. Common pitfalls to avoid. Do not run multiple pipelines in the same directory. The framework uses temporary files during execution, and overlapping runs can corrupt each other's data. If you need parallel execution, use separate directories and run each with a different working directory flag. Do not store sensitive data like passwords in your config.json file. Use environment variables instead. The framework reads variables prefixed with AGE_ automatically. Set them in your shell profile or a .env file, and the framework handles the rest.

Another frequent mistake is ignoring timezone settings. Age Of The Exploration defaults to UTC for all timestamps. If your data sources use local time, convert them in your transform stage using the timezone filter. Failing to do this causes off-by-one-day errors in daily reports, which is a silent bug that propagates into downstream analytics. I have seen this trip up finance teams during quarterly reconciliations. When Age Of The Exploration is not the right tool. The framework excels at batch processing of structured data. It struggles with unstructured data like images or natural language text. If your project involves those data types, consider pairing it with a specialized tool or using the custom script extension. The extension allows you to write Python code for any transformation that the built-in operations cannot handle. However, custom scripts bypass the built-in validation layer, so test them thoroughly before deploying. A malformed custom script can cause the framework to skip validation entirely and produce corrupted output without any warnings. There is also a hard limit on dataset size for the built-in memory processor. If your individual files exceed 2 gigabytes, the framework will load them entirely into RAM and may crash your system. In those cases, use the streaming mode by adding stream: true to your input source configuration. This processes data in chunks and keeps memory usage under 500 megabytes regardless of file size. The tradeoff is that streaming mode disables certain operations like global aggregations. If you need aggregations on large datasets, write the intermediate results to disk and run a second pipeline pass over them.
I have found that the best approach for enterprise-scale projects is a hybrid setup. Use Age Of The Exploration for the standard pipeline work and integrate it with a separate tool like Apache Airflow for orchestration and monitoring. The framework provides an HTTP endpoint at localhost:8080 that Airflow can call to trigger runs. This gives you the simplicity of Age Of The Exploration with the robustness of a dedicated orchestrator. It adds about an hour of setup time, but it pays off within the first month of operation.