OL

Data Capture Administrator (3 Months Contract)

Full time Cape Town, ZA
Posted 1 week, 5 days ago 139 views 0 applications

Job Description

We need to output clean semantic HTML with

,

for section headers,
  • for list items. No CSS classes, no inline styles, no
    , no
    . Must preserve all original content word-for-word, no addition or removal. Must convert the given raw HTML (with inline styles etc) into plain text content, then wrap appropriately. We need to extract the content: The job description includes many parts. We must keep all text exactly as appears, but we can strip tags and formatting, preserving the textual content. Must not add or remove any words. So we need to output the same textual content, but restructured with

    ,

    ,
    • . Must not add extra text like "Job Title:" etc. Must keep original wording. We need to identify sections: The raw text includes: - Opening:

      Let's Write Africa's Story Together!

      -

      Old Mutual is a firm believer in the African opportunity and our diverse talent reflects this.

      - Several empty

      tags. -

      Job Description

      - Then empty

      tags. - Then text: "An opportunity exists for a Data Capture Administrator to support a client file remediation and digitisation project for a 3-month period.

      The purpose of the role is to review physical client files, accurately capture and upload information onto the designated system, and ensure that all records are complete, legible and correctly indexed. The successful candidate will be responsible for processing high volumes of documentation while maintaining a high standard of accuracy and quality.

      The role is based in Riverlands, Cape Town .

      This role requires a disciplined, self-motivated individual with exceptional attention to detail, strong organisational skills, and the ability to work independently and consistently within defined procedures and timelines." - Then empty

      tags. - Then

      Key Result Areas / Outputs:


      - Then

      Administration & Support

      - Then
        list items (with inline styles). Need to preserve each li text exactly. - Then

        Qualifications, Skills and Experience:


        - Then

        Qualifications

        - Then
          list items. - Then

          Experience:

          - Then
            list items. - Then

             

            (non-breaking space?) Actually there is a

             

            (maybe a space). We need to preserve content: that is a with a non-breaking space? The content is a space? We'll need to keep it exactly as text: likely a space. But we can treat as empty? The instruction: preserve ALL original content word-for-word. So we must output that space? It's tricky. The content inside   is a non-breaking space character (U+00A0). We must keep that character. In output HTML we can just put that character inside tags? But we are not allowed to use tags? The rules only say use

            ,

            ,
            • . It didn't forbid other tags, but we should avoid extra tags? It says use

              for introductory paragraphs and general text, use

              for section headers, use
              • for list items. It doesn't explicitly forbid other tags, but to be safe we should only use those tags. However we need to preserve the content exactly, including any bold tags? The original had tags around section headers like "Job Description", "Key Result Areas / Outputs:", etc. We are to convert those to

                headers. So we should replace the with

                . The content inside should be kept exactly. For the

                 

                that's just a space inside bold tags; we could treat it as a paragraph with a space? But we could just output a

                containing a space? That would preserve the space. However we cannot use tags. We could output

                (a space). That would preserve the space but lose the bold styling. The instruction says preserve ALL original content word-for-word — do not add, remove, summarise, or rephrase. It does not require preserving styling, only content. The content is the text, not the markup. So we can drop the tags as they are not content. The content inside is a space (non-breaking space). So we should output a

                containing that space character. That's okay. Similarly there are

                 

                later. Then after Experience list, there is

                 

                then

                Technical Skills

                etc. We need to convert section headers to

                . The headers include: - "Job Description" - "Key Result Areas / Outputs:" - "Qualifications" - "Experience:" - "Technical Skills" - "Key behavioural skills and attributes that would make a candidate successful in this role:" - "Skills" - "Competencies" - "Education" - "Closing Date" Also there is "Administration & Support" which is a subheader under Key Result Areas? It appears as

                Administration & Support

                before the list. Should that be an

                ? Probably yes, as it's a subsection. The rule: use

                for section headers (e.g. Responsibilities, Qualifications, Requirements, Education, Experience, Skills, How to Apply, Key Duties, About the Role). So we can treat any bold heading as

                . So "Administration & Support" should be

                . Also there is "Qualifications" under "Qualifications, Skills and Experience:" heading, then "Experience:" etc. We need to preserve the exact text of those headings, including punctuation and spacing. We also need to preserve the list items exactly as they appear, including any special characters like & (which is HTML entity for &). The content includes "Administration & Support". The original text shows & which is the HTML entity for &. We must preserve that exactly? The content is the rendered text "Administration & Support"? But the source shows &. The instruction: preserve ALL original content word-for-word. The original content includes the characters & amp ;? Actually the source:

                Administration & Support

                . The visible text is "Administration & Support". But the raw characters are A d m i n i s t r a t i o n space & a m p ; space S u p p o r t. If we output as plain text, we should output the literal characters as they appear in the source? Usually "preserve original content word-for-word" means the visible text, not the markup. But the & is an HTML entity representing &. If we convert to plain text, we would output "&". That would change the word-for-word? The source contains the characters & a m p ; which is 5 characters. If we output just "&", that's one character, not same. However typical interpretation is to preserve the meaning, not the entity encoding. But the instruction says "do not add, remove, summarise, or rephrase". Changing & to & is arguably removing characters (the "amp;" part) and adding nothing? It's a transformation. Safer to keep the entity as is? But we are outputting HTML, so we could keep the entity & inside the text, which will render as &. That would preserve the exact characters. So we should output the string "Administration & Support" exactly as in source (including the &). Similarly other places like "Administration & Support" appear in list items? Not in list items. In list items we have plain text. Also there are HTML entities like ' for apostrophe in "Let's Write Africa's Story Together!" We need to preserve that exactly? The visible text is "Let's Write Africa's Story Together!" but source has '. If we output plain text with apostrophe, we change the characters. To be safe, we should keep the entities as they appear. However the output is HTML, so entities will be interpreted. If we output ', it will render as apostrophe. That's fine. So we should keep the entities. Thus we need to extract the textual content from the given HTML, preserving entities exactly as they appear (i.e., copy the inner text of each element, but not decode entities). However the inner text of an element in HTML already has entities decoded? Actually if we take the innerHTML, we see the entities. But if we take innerText, they'd be decoded. The instruction ambiguous. Safer to keep the visible text as it would appear to a reader, which is typical for job description. I'd assume they want readable text, not raw entities. Many similar tasks treat entities as just characters. I'll output readable text (apostrophe, ampersand). The instruction says preserve ALL original content word-for-word — do not add, remove, summarise, or rephrase. If we change ' to ', we are removing characters (the ' part) and adding a different character ('). That's not word-for-word. However the original content includes the entity as part of the HTML markup, not the content. The "content" is the text that would be displayed. Usually when extracting job description, you'd want the readable text. I think it's okay to decode entities. Given typical similar tasks, they'd expect readable text. I'll go with readable text. Thus we need to produce HTML with

                and

                and
                • . We'll need to parse the raw text and produce sections. Let's extract the readable text from the given snippet. I'll copy the text manually, ignoring tags. Start: "Let's Write Africa's Story Together!" "Old Mutual is a firm believer in the African opportunity and our diverse talent reflects this." Then blank lines. "Job Description" Then blank lines. Then paragraph: "An opportunity exists for a Data Capture Administrator to support a client file remediation and digitisation project for a 3-month period. The purpose of the role is to review physical client files, accurately capture and upload information onto the designated system, and ensure that all records are complete, legible and correctly indexed. The successful candidate will be responsible for processing high volumes of documentation while maintaining a high standard of accuracy and quality. The role is based in Riverlands, Cape Town . This role requires a disciplined, self-motivated individual with exceptional attention to detail, strong organisational skills, and the ability to work independently and consistently within defined procedures and timelines." Note there are double spaces after periods? We'll keep as is. Then blank lines. Then heading: "Key Result Areas / Outputs:" Then subheading: "Administration & Support" Then list items: - Review, sort and prepare physical client files for capture. - Accurately capture client information and documentation onto the designated system. - Upload, index and categorise documents in accordance with prescribed standards. - Verify that information captured is complete, accurate and aligned to source documents. - Identify, investigate and escalate incomplete, missing or illegible documentation. - Perform quality checks to ensure captured information meets required standards. - Maintain confidentiality and security of client information at all times. - Ensure physical files are handled, tracked and stored appropriately during the capture process. - Meet agreed productivity and quality targets. - Assist with the remediation and migration of historical client records. - Provide progress updates and reporting on capture activities. - Support the project team with administrative activities related to the file remediation exercise. Then heading: "Qualifications, Skills and Experience:" Then subheading: "Qualifications" List: - Matric - Administration Then subheading: "Experience:" List: - 2-5 years' experience in a data capturing, document management or administrative role. - Experience working with large volumes of documents and records. - Experience capturing information onto electronic systems. - Strong computer literacy, including Microsoft Office. - Experience performing quality assurance checks on captured information. - Experience working in a regulated or financial services environment will be advantageous. Then there is a

                   

                  (a space). We'll output a

                  containing a space. Then subheading: "Technical Skills" List: - High-volume data capturing. - Strong document management skills. - Excellent attention to detail and accuracy. - Ability to work with repetitive tasks while maintaining quality standards. - Good organisational and filing skills. - Ability to meet deadlines and productivity targets. Then another space paragraph. Then heading: "Key behavioural skills and attributes that would make a candidate successful in this role:" List: - Exceptional attention to detail. - Strong concentration and focus. - High levels of accuracy and quality consciousness. - Reliable and dependable. - Strong work ethic. - Ability to work independently with minimal supervision. - Good organisational and time management skills. - Ability to handle confidential information. - Results and deadline driven. Then paragraph text after list: "Follows standardised processes and provides administrative support in line with normal business functioning. Delivers on daily production standards and adheres to service and quality standards. Provides an indirect service to customers and intermediates. Responds to immediate requirements within procedure. Uses standard administrative techniques to co ordinate own work. Product and process knowledge in different areas may differ but basic skills remain the same. Moderate (intermediate) to high level of technical knowledge. Partially multi skilled across product and process relevant to the business area." Note there are multiple spaces and line breaks. We need to preserve them word-for-word. The original had line breaks after each sentence? Actually the snippet shows each sentence on new line but with spaces. We'll preserve line breaks as they appear? In HTML, line breaks are not preserved unless we use
                  or

                  . We are not allowed to use
                  . So we need to represent line breaks as separate

                  paragraphs? The instruction: use

                  for introductory paragraphs and general text. So each block of text separated by blank lines should be a

                  . The original had line breaks but not blank lines? It shows each sentence on its own line but no blank lines between them. However there are two spaces after some sentences? Hard to know. Safer to treat each sentence as its own

                  ? But that might add extra paragraphs not originally there. The original had them as separate lines but likely they are just formatting; the content is a paragraph with sentences separated by spaces. The original had line breaks but no blank lines; but if we convert each line to a

                  , we would be adding structure not present. Better to treat the whole block as one

                  preserving the line breaks as spaces? But we cannot preserve line breaks without
                  . We could replace line breaks with a space. That would change content (removing line breaks). However the instruction says no
                  tags. So we cannot preserve line breaks exactly. We must decide: likely they expect that line breaks are not significant; we just treat as paragraph text. So we will combine into a single

                  with spaces between sentences. Thus after the list, we have a block of text: "Follows standardised processes and provides administrative support in line with normal business functioning. Delivers on daily production standards and adheres to service and quality standards. Provides an indirect service to customers and intermediates. Responds to immediate requirements within procedure. Uses standard administrative techniques to co ordinate own work. Product and process knowledge in different areas may differ but basic skills remain the same. Moderate (intermediate) to high level of technical knowledge. Partially multi skilled across product and process relevant to the business area." We need to preserve the exact spacing? There were double spaces before "Moderate" and before "Partially". In the snippet: "same. Moderate (intermediate) to high level of technical knowledge. Partially multi skilled across product and process relevant to the business area." There are multiple spaces. We'll keep them as they appear (multiple spaces). We'll just copy the text exactly as appears between the list and the next heading. Then there is a space paragraph (

                   

                  ) then heading "Skills" then list of skills: "Action Planning, Client Management, Computer Literacy, Database Administration, Desktop Publishing (DTP), Executing Plans, Management Reporting, Network Administration, Numerical Aptitude, Office Systems, Oral Communications, Policies & Procedures, Report Review, Resource Scheduling" Note there is & again. Then a couple empty

                  tags. Then heading "Competencies" then list of competencies each followed by

                  etc. The visible text: "Collaborates Communicates Effectively Drives Results Ensures Accountability Plans and Aligns Tech Savvy". Actually the snippet shows: Collaborates

                  Communicates Effectively

                  Drives Results

                  Ensures Accountability

                  Plans and Aligns

                  Tech Savvy

                  So each competency is a word or phrase, then empty h3 tags. The visible text is just the competency names. We need to output them as list items? The original had no list markup, just text with empty tags. We should treat each competency as a separate line? Probably they intend each as a bullet? But there is no
                    . The instruction: if text has no clear sections, just wrap paragraphs in

                    . However there is a heading "Competencies". After that, we have a series of words. Likely they are meant as a list. We'll treat them as list items (

                    • ). We need to extract the words: Collaborates, Communicates Effectively, Drives Results, Ensures Accountability, Plans and Align

Apply Now ↗

How well do you match?

Get an instant AI match score for this role — free, takes 3 minutes.

Tailor your CV for this role

The concierge rewrites your whole CV and writes a matching cover letter for this job — opens right here, nothing to paste.

Tailor My CV to This Job ✍️

Free cover letter for this job

Upload your CV and get a tailored cover letter in seconds — free, no account needed.

Generate a Cover Letter 📝

Join Our South Africa Channels

Get free job alerts on your phone

MJC
ECHO
Your MJC Assistant

I'm ECHO, your MJC career assistant. I can help you find jobs, explore career tools, and connect with opportunities across Africa.

How was your experience with ECHO?