AI Research & Evaluation Specialist - 24-MAG (Remote)

24-MAG is looking for detail-oriented professionals and postgraduate students to join their AI research initiative. This worldwide remote role focuses on benchmarking and improving cutting-edge language models through structured research, evaluation, and question development. It is a fantastic opportunity for those with strong analytical skills to contribute directly to the advancement of AI technology while working in a flexible, asynchronous environment.

Remote AI Jobs Feed
Role Description

We are sharing a specialised part-time consulting opportunity for detail-oriented professionals and postgraduate students with strong research, reasoning, communication, and information-synthesis skills. This role supports an advanced AI research initiative focused on benchmarking and improving cutting-edge language models. Selected professionals will create high-quality research questions, develop accurate reference answers, review information from multiple sources, and prepare structured evaluation materials used to assess and improve AI systems.

### Key Responsibilities

* **Research & Information Analysis**

+ Conduct structured online research across a variety of STEM and non-STEM topics
+ Gather relevant information from multiple sources and synthesise findings clearly
+ Evaluate source quality, relevance, and consistency
+ Identify gaps, contradictions, and important contextual details within research materials
* **Question & Answer Development**

+ Create high-quality research questions designed to test advanced AI systems
+ Develop accurate, well-supported reference answers
+ Structure questions and solutions so they can be evaluated consistently
+ Apply strong reasoning and written communication across interdisciplinary subject areas
* **AI Evaluation Materials**

+ Develop structured materials used to benchmark and assess language-model performance
+ Review AI-generated responses for accuracy, reasoning quality, completeness, and clarity
+ Identify subtle errors, omissions, and unsupported conclusions
+ Provide clear written feedback that supports model improvement
* **Quality & Independent Work**

+ Follow project instructions and evaluation criteria carefully
+ Maintain strong accuracy across independent research and review tasks
+ Manage assignments within a flexible asynchronous workflow
+ Incorporate feedback and project guidance into subsequent work

Qualifications

* Approximately 2–3 years of relevant professional experience in a STEM, non-STEM, research, analytical, or related field
* Current or recently completed master's-level studies
* Alternatively, a bachelor's degree combined with relevant professional experience
* Strong online research and information-synthesis skills
* Excellent written communication and analytical reasoning
* Ability to gather information from diverse sources and summarise it clearly
* Strong attention to detail and comfort working independently
* Availability to contribute approximately 10–20 hours per week
* Access to a compatible MacBook throughout the project

Requirements

* A MacBook with an Apple Silicon M-series processor is required
* Compatible devices include M-series MacBook Pro and MacBook Air models
* The device must support the project's required macOS environment, including macOS 15 or higher where applicable
* A desktop or non-Apple device cannot substitute for the required MacBook setup

Nice to Have

* An interdisciplinary academic or professional background
* Previous research-assistant, analyst, data, or evaluation experience
* Familiarity with AI tools and large language models
* Experience developing research questions, reference answers, or structured assessment materials
* Background in fact-checking, content review, quality assurance, or data annotation
* Experience working independently on research-intensive projects

Why This Opportunity

* Contribute directly to research used to benchmark advanced AI systems
* Work across a broad range of interdisciplinary topics
* Apply research, reasoning, and writing skills to practical AI evaluation tasks
* Build experience creating high-quality training and evaluation materials
* Participate in flexible part-time work with an immediate-start preference

Contract Details

* Independent contractor role
* Fully remote with flexible scheduling
* Expected commitment of approximately 10–20 hours per week
* Immediate availability is preferred
* Competitive rates between $45–$55 per hour depending on experience and project scope
* MacBook with a compatible Apple Silicon M-series chip is required
* The application process includes resume review and a brief approximately 15-minute AI interview focused on research and reasoning skills
* Selected applicants will complete a brief paid assessment
* Weekly payments via Stripe or Wise
* Projects may be extended, shortened, or adjusted depending on project requirements and performance
* Work will not involve access to confidential or proprietary information from any employer, client, or institution.