Discover what should be included in an academic journal archive, from published articles and metadata to submissions, peer reviews, editorial records, users, and files.
TL;DR
- A complete academic journal archive should contain more than published article PDFs.
- Journal identity, volumes, issues, publications, and metadata form its foundation.
- Submissions, manuscript versions, peer reviews, and editorial records preserve publishing history.
- Author, reviewer, editor, and user information should be properly organized.
- Files, documentation, and integrity information make archives more portable and reliable.
Academic journals build a valuable publishing record with every manuscript they receive, review, revise, and publish.
That record includes much more than the articles eventually made available to readers. Manuscript versions, peer-review activity, editorial decisions, contributor information, publication metadata, and associated files can all become important records over the life of a journal.
A comprehensive academic journal archive brings these records together in an organized and accessible format. This can help publishers preserve their history, maintain institutional continuity, support future platform changes, and keep important publishing information available beyond the day-to-day operation of their journal management system.
Alkademy Press takes this broader view of journal archiving through its Press Archive Package (PAP). The package provides a structured export of a journal’s publishing records and associated files, helping publishers retain information across the journal’s publishing lifecycle. In this guide, we look at what a complete academic journal archive should contain and the considerations publishers should make when preserving their journal data.
Recommended for You: Press Archive Package (PAP): The Complete Guide to Academic Journal Archives
What Is an Academic Journal Archive?
An academic journal archive is a structured collection of records and files that preserves information about a journal’s publishing activity over time.
That can include the public publishing record such as articles, issues, and metadata as well as restricted records documenting editorial activity.
A useful archive may contain:
- Journal identity and configuration
- Volumes and issues
- Published articles and publication metadata
- Manuscript submissions
- Manuscript versions and revisions
- Peer-review records, where retention is appropriate
- Editorial decisions and relevant history
- Authors, editors, reviewers, and other contributors
- Associated manuscript and publishing files
- Archive documentation
- File-integrity information
Not every journal needs every category. The appropriate scope should be determined before export rather than assuming that every field in a publishing platform belongs in the permanent archive.
Why Academic Journals Need Comprehensive Archives
Journals can operate for many years, often through changes in editorial teams, institutions, publishing workflows, and technology platforms. Without a well-organized archive, important historical information can become difficult to locate or may remain tied to a particular system.
A comprehensive archive can help with:
- Long-term preservation: Keep important publishing records available over time.
- Platform migration: Provide a reference when moving journal data to another system.
- Institutional continuity: Help new editorial or administrative teams understand the journal’s history.
- Record keeping: Retain information about submissions, reviews, revisions, and publications.
- Data portability: Maintain a structured copy of journal data outside the live publishing environment.
For publishers, this means thinking about archiving as part of responsible journal management rather than something to consider only when a platform change becomes necessary.
Essential Things to Include in an Academic Journal Archive
A complete academic journal archive should capture the information that gives context to a journal’s publishing activities. While every journal may have different requirements, these ten categories provide a strong foundation for preserving its publishing record.
Journal Identity and Core Information
Start with the information that defines the journal itself. This should include details such as:
- Journal name and description
- ISSN
- Journal type
- Website information
- Journal settings
- Submission status
- Article Processing Charge settings
- Branding assets
Preserving this information helps ensure that the archive remains identifiable and useful even when separated from the original publishing platform.
Volumes and Issues
The archive should preserve the journal’s publication structure, including:
- Volume numbers and publication years
- Issue numbers and titles
- Issue descriptions
- Publication dates
- Issue status
- Cover images where applicable
This makes it easier to reconstruct the journal’s published catalog and understand how its content was organized.
Published Articles and Publication Metadata
Published articles are central to any journal archive, but the article files should be accompanied by relevant metadata.
Important information can include:
- Article title
- DOI
- Publication date
- Page information
- Issue and volume relationship
- Publication status
Structured metadata makes published content easier to identify, organize, and preserve.
Manuscript Submission Records
A journal archive should also preserve information about manuscripts submitted for consideration, not only those that were published.
Submission records can include:
- Manuscript title
- Abstract
- Keywords
- Author information
- Submission status
- DOI where applicable
- Associated metadata
- Creation and update dates
These records provide useful context around the journal’s editorial activity.
Manuscript Versions and Revisions
Manuscripts can change significantly during peer review and editorial revision. Where version records are retained, the archive should preserve information that distinguishes each version.
This can include:
- Version number
- Review round
- Filename
- File size
- Associated metadata
- Location of the manuscript file
Keeping these records helps maintain a clearer history of how submissions developed.
Peer-Review Records
Peer review is an important part of scholarly publishing, so relevant review records should be considered when defining the scope of an archive.
These may include:
- Review rounds
- Reviewer references
- Ratings
- Decisions
- Feedback
- Review status
- Due dates
- Conflict-of-interest information
- Review timestamps
The exact information preserved should reflect the journal’s review model and privacy policies.
Editorial History and Comments
Editorial activity can provide important context that is not visible in the final published article.
An archive may preserve:
- Editorial comments
- Review comments
- Submission-related discussions
- Comments associated with manuscript versions
- Editorial timestamps
- Decision-related records
Preserving this history can be particularly useful for institutional record keeping and understanding the development of a manuscript.
Authors, Co-Authors, Editors, and Reviewers
People and their relationships to journal records should be preserved in a structured way.
Relevant information can include:
- Names
- Affiliations
- ORCID
- Journal roles
- Author order
- Submission relationships
- Contributor status
Keeping these relationships intact helps connect people to the manuscripts, reviews, and publications with which they were involved.
Manuscript and Journal Files
An archive should account for the actual files associated with the journal, rather than preserving metadata alone.
Depending on the journal, these can include:
- Manuscript files
- Journal logo
- Banner and branding assets
- Other associated publishing files
File references should also be connected to the records they belong to, making the archive easier to understand and use.
Archive Documentation and Integrity Information
Finally, an archive should explain what it contains and provide information that helps publishers verify its contents.
Useful archive documentation can include:
- Archive format and version
- Export date and time
- Journal identification
- Export options
- Record counts
- Export warnings
- File integrity checksums
For example, the Press Archive Package (PAP) from Alkademy Press includes a manifest.json file for archive information and a checksums.sha256 file for file integrity verification.
What Should Not Be Included in an Academic Journal Archive?
A comprehensive archive does not mean exporting every piece of information stored by a publishing system. Publishers should consider security, privacy, and the purpose of the archive when determining what belongs in an export.
Passwords and Authentication Credentials
Passwords and authentication credentials should not be part of a journal archive. They provide access to accounts rather than contributing to the journal’s publishing history and can create unnecessary security risks if exported.
In the Press Archive Package, passwords are excluded entirely.
Unnecessary Sensitive Information
Publishers should avoid including personal or sensitive information that is not needed for the archive’s intended purpose.
Before exporting, consider:
- Whether each data field serves a legitimate archival purpose
- Whether personal information needs to be retained
- Who will have access to the archive
- How long the information needs to be preserved
A useful archive should be comprehensive without collecting information simply because it happens to exist in the system.
Reviewer Information That Could Compromise Anonymity
Reviewer information requires particular care when a journal uses anonymous or blinded peer review.
Review records may contain reviewer references, feedback, decisions, or conflict-of-interest information. Publishers should ensure that archived data is handled consistently with their review policies and applicable privacy requirements.
The objective is to preserve useful editorial records without unintentionally exposing information that should remain confidential.
Academic Journal Archive vs. Published Article Archive
A published-article archive and a comprehensive journal archive have different scopes.
| Published Article Archive | Comprehensive Journal Archive |
| Published article files | Published article files |
| Basic issue information | Journal identity and configuration |
| Publication metadata | Detailed publication metadata |
| Final published versions | Submissions and manuscript versions where retained |
| Basic contributor information | Contributor roles and relationships |
| Usually no editorial history | Relevant editorial and decision records |
| Limited associated files | Manuscript, supplementary, and journal files |
| Limited documentation | Manifest, export details, warnings, and integrity information |
If the goal is simply to preserve what readers can access, an article archive may be sufficient.
If the goal is institutional continuity, migration, editorial record keeping, or preservation of the journal’s publishing history, a broader archive is more appropriate.
How to Organize an Academic Journal Archive

A well-organized archive should make it easy to identify different types of information and understand how they relate to one another. Separating records into logical categories also makes the archive easier to maintain, review, and potentially migrate.
Journal Information
Keep information that identifies and describes the journal together, including its name, ISSN, description, settings, and branding.
Published Catalog
Organize volumes, issues, publications, DOIs, publication dates, and related metadata so the journal’s published structure can be understood independently.
Editorial Records
Keep submissions, manuscript versions, review rounds, reviews, decisions, and editorial comments together or clearly linked through their relevant identifiers.
Users and Contributors
Preserve authors, co-authors, editors, reviewers, and other users alongside the relationships that connect them to journal activities.
Associated Files
Store manuscript and journal files in a predictable structure and maintain clear references between files and the records they belong to.
Archive Documentation
Include information that explains the archive itself, such as its format, version, export details, record counts, warnings, and file integrity information.
Best Practices for Academic Journal Data Preservation
Good journal data preservation requires more than simply creating a copy of existing records. Publishers should consider how the information will remain understandable, accessible, and trustworthy over time.
- Preserve structured metadata: Keep important information with the records it describes.
- Retain manuscript versions: Where applicable, preserve relevant versions and their relationships to review rounds.
- Keep editorial history: Retain appropriate review and editorial records that document the publishing process.
- Verify file integrity: Use checksums or other verification methods to identify unexpected file changes.
- Protect personal information: Apply appropriate privacy controls to author, reviewer, and user data.
- Control archive access: Limit access to authorized people and handle exported archives securely.
- Use durable storage: Store archives in reliable locations and maintain appropriate copies for long-term access.
- Document the archive: Make it clear what the archive contains, when it was created, and whether any records or files could not be included.
The objective is to create an archive that remains useful not only when it is created, but also when publishers need to access, verify, or use it years later.
What Makes an Academic Journal Archive Complete?
There is no single archive structure that works for every journal. The right scope depends on the journal’s publishing model, policies, and record-keeping requirements.
However, a comprehensive archive should generally answer four basic questions:
- What is the journal?
Its identity, settings, catalog, and branding. - What has the journal published?
Its volumes, issues, articles, publication metadata, and associated files. - How did the journal’s content move through publishing?
Its submissions, revisions, peer reviews, editorial activity, and decisions. - Who and what are connected to those records?
Its authors, co-authors, editors, reviewers, users, metadata, and files.
This broader approach provides a much clearer picture of a journal’s publishing history and creates a stronger foundation for preservation and data portability.
How Alkademy Press Supports Academic Journal Archiving
A complete journal archive needs to bring different types of publishing information together in a way that remains organized and understandable. Alkademy Press addresses this through its Press Archive Package (PAP), a structured .zip export containing the records and files associated with a single journal.
PAP can bring together:
- Journal information: Identity, settings, members, and branding references
- Published catalog: Volumes, issues, publications, DOIs, and publication details
- Editorial records: Submissions, manuscript versions, review rounds, reviews, and comments
- Users and contributors: Authors, co-authors, editors, reviewers, and other users
- Associated files: Manuscripts and journal branding files
- Archive information: Export details, record counts, warnings, and file checksums
This approach gives publishers a structured representation of their journal’s publishing history rather than a collection of disconnected files.
What Should Publishers Ask an Archive or Export Provider?
Before choosing an archiving solution, publishers should ask specific technical questions rather than relying on a claim that the platform “exports your data.”
Records
- Which journal records are exported?
- Are submissions included?
- Are rejected and withdrawn submissions included?
- Are manuscript versions retained?
- Are review rounds and editorial decisions represented?
Metadata
- Which metadata fields are exported?
- Are identifiers preserved?
- Are relationships between records maintained?
- Is the schema documented?
Files
- Which file types are included?
- Are supplementary files included?
- How are files linked to their records?
- Are checksums generated?
Documentation
- Is the archive format documented?
- Is there a manifest?
- Are record counts provided?
- Are export errors and exclusions reported?
Security and privacy
- Are passwords and credentials excluded?
- How is confidential peer-review information handled?
- Can archive access be restricted?
- What personal information is retained?
Portability
- Is the archive readable independently of the live platform?
- Can the data be imported into another system?
- Is the format documented well enough for another developer or publisher to interpret?
These questions help distinguish a genuine archival export from a simple collection of downloadable files.
Academic Journal Archive Checklist
Use this checklist when reviewing an existing journal archive or evaluating an archiving approach:
- Journal identity, ISSN, and settings
- Volumes and publication years
- Issues and issue metadata
- Published articles and publication metadata
- DOI information
- Manuscript submission records
- Manuscript versions and revisions
- Peer-review records and review rounds
- Editorial history and comments
- Authors and co-authors
- Editors, reviewers, and other journal members
- User records
- Manuscript files
- Journal branding files
- Archive manifest and documentation
- Record counts and export warnings
- File integrity information
- Appropriate privacy and access controls
The checklist can help publishers identify gaps before creating, transferring, or reviewing a journal archive.
FAQs
A comprehensive archive should include the journal’s identity and settings, volumes, issues, publications, metadata, submissions, manuscript versions, peer-review records, editorial history, contributors, users, and associated files. Archive documentation and integrity information are also useful for understanding and verifying the exported records.
It depends on the journal’s record-retention policies and the purpose of the archive. If rejected submissions form part of the journal’s historical or editorial records, publishers may choose to preserve them. Any archived information should be handled according to applicable privacy, confidentiality, and retention requirements.
Relevant manuscript versions should be stored with information that identifies the submission, version number, review round, filename, and associated metadata. Keeping the relationship between each version and its manuscript record makes the editorial history easier to understand and preserve.
A journal backup is generally intended to help restore a system or recover data after a failure. A journal archive is designed to preserve publishing records for future access, reference, portability, or migration. An archive can therefore provide a structured record of the journal’s history rather than simply serving as a recovery copy.
Final Thought On What Should Be Included in an Academic Journal Archive
A journal’s history extends well beyond the articles that appear on its website. Submissions, manuscript revisions, peer reviews, editorial activity, contributor information, metadata, and associated files can all contribute to a complete record of the journal’s publishing work.
A thoughtful academic journal archive preserves these elements in an organized way, giving publishers a stronger foundation for long-term preservation, data portability, institutional continuity, and future platform changes.
With the Press Archive Package, Alkademy Press provides a structured way to bring these records and files together in a portable archive.
Want greater control over your journal’s publishing records? Explore Alkademy Press to manage your journal’s publishing workflow and preserve its data with a structured, portable archive.

