Publications

2026

  1. ASE’26
    ACM Artifacts Available ACM Artifacts Evaluated: Functional

    ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories

    In Proceedings of the 41st IEEE/ACM International Conference on Automated Software Engineering (ASE’26), October 12-16, 2026, Munich, Germany 4 citations 10 stars
    Abstract
    Most repository-level code translation and validation techniques have been evaluated on a single source-target programming language (PL) pair, owing to the complex engineering effort required to adapt new PL pairs. Programming agents can enable PL-agnosticism in repository-level code translation and validation: they can synthesize code across many PLs and autonomously use existing tools specific to each PL’s analysis. However, state-of-the-art has yet to offer a fully autonomous agentic approach for repository-level code translation and validation of large-scale programs. This paper proposes ReCodeAgent, an autonomous multi-agent approach for language-agnostic repository-level code translation and validation. Users only need to provide the project in the source PL and specify the target PL for ReCodeAgent to automatically translate and validate the entire repository. ReCodeAgent is the first technique to achieve high translation success rates across many PLs. We compare the effectiveness of ReCodeAgent with four alternative neuro-symbolic and agentic approaches to translate 118 real-world projects, with 1,975 LoC and 43 translation units for each project, on average. The projects cover 6 PLs and 4 PL pairs. Our results demonstrate that ReCodeAgent consistently outperforms prior techniques on translation correctness, improving test pass rate by 60.8% on ground-truth tests, with an average cost of $15.3. We also perform process-centric analysis of ReCodeAgent trajectories to confirm its procedural efficiency. Finally, we investigate how the design choices (a multi-agent vs. single-agent architecture) influence ReCodeAgent performance: on average, the test pass rate drops by 40.4%, and trajectories become 28% longer and persistently inefficient.
    BibTeX
    @article{ibrahimzada2026recodeagent,
      title={ReCodeAgent: A Multi-Agent Workflow for Language-agnostic Translation and Validation of Large-scale Repositories},
      author={Ibrahimzada, Ali Reza and Paulsen, Brandon and Kroening, Daniel and Jabbarvand, Reyhaneh},
      journal={arXiv preprint arXiv:2604.07341},
      year={2026}
    }
  2. PHD THESIS

    Neuro-Symbolic Code Translation and Validation

    Ali Reza Ibrahimzada
    PhD Thesis, University of Illinois Urbana-Champaign
    Abstract
    Software systems often outlive the programming languages (PLs) in which they were first written. As hardware platforms, deployment environments, security expectations, and developer ecosystems evolve, organizations must repeatedly migrate valuable code from legacy or ill-suited languages to modern ones. Automated code translation promises to reduce this migration effort, but reliable repository-level translation remains difficult. Traditional transpilers encode useful language knowledge in their design, yet they require substantial language-specific engineering and often produce brittle, non-idiomatic code. Large language models (LLMs) produce more natural and idiomatic translations, yet they struggle on real repositories, where correctness depends on cross-file dependencies, library APIs, type systems, and developer-written tests. This dissertation presents work on neuro-symbolic and agentic code translation and validation. To make repository-level translation more practical, this dissertation develops techniques that combine the generative abilities of LLMs with program analysis, execution feedback, test-based validation, and multi-agent workflows for translation and validation. These techniques expose the structure of real codebases, guide translation decisions with semantic context, and use validation not merely as a final check but as a source of feedback for detecting and repairing translation failures. This dissertation presents contributions along two key directions. The first direction studies why LLM-based code translation fails in realistic settings. Although LLMs can often translate small, self-contained examples, real repositories require global reasoning about language-specific features, call chains, libraries, and tests. This dissertation presents an empirical study of these failures and develops a taxonomy of translation bugs. The taxonomy establishes a foundational understanding of LLM translation limitations and serves as a guide for designing practical tools to mitigate the most common failures. The second direction develops techniques for translating, validating, and repairing repository-level code. It first presents AlphaTrans, a neuro-symbolic pipeline that decomposes a repository into smaller fragments, translates them in dependency-aware order, and validates partial translations in isolation while preserving the behavior exercised by existing tests. It then presents MatchFixAgent, a language-agnostic validation and repair framework that combines approximate semantic analyses with specialized agents for test generation and repair. Finally, it presents ReCodeAgent, an end-to-end multi-agent workflow that assigns analysis, planning, translation, and validation to dedicated agent scaffolds; ablation studies show that this division of labor is necessary, as monolithic single-agent instantiations cannot achieve comparable effectiveness. Together, these contributions show that LLMs alone are not sufficient for dependable code translation, but can become effective components of larger systems when guided by program analysis and subjected to rigorous validation. The techniques in this dissertation move automated code translation from simple examples toward realistic software modernization, helping developers understand translation failures, produce more reliable translations, and reason more systematically about the correctness of translated code.
  3. ICML’26

    MatchFixAgent: Language-Agnostic Autonomous Repository-Level Code Translation Validation and Repair

    In Proceedings of the 43rd International Conference on Machine Learning (ICML’26), July 6-11, 2026, Seoul, South Korea 11 citations 2 stars
    Abstract
    Code translation transforms source code from one programming language (PL) to another. Validating the functional equivalence of translation and repairing, if necessary, are critical steps in code translation. Existing automated validation and repair approaches struggle to generalize to many PLs due to high engineering overhead, and they rely on existing and often inadequate test suites, which results in false claims of equivalence and ineffective translation repair. To bridge this gap, we develop MatchFixAgent, a large language model (LLM)-based, PL-agnostic framework for equivalence validation and repair of translations. MatchFixAgent features a multi-agent architecture that divides equivalence validation into several sub-tasks to ensure thorough and consistent semantic analysis of the translation. We compare MatchFixAgent’s validation and repair results with four repository-level code translation techniques. Our results demonstrate that MatchFixAgent produces (in)equivalence verdicts for 99.2% of translation pairs, with the same equivalence validation result as prior work on 72.8% of them. When MatchFixAgent’s result disagrees with prior work, we find that 60.7% of the time MatchFixAgent’s result is actually correct. In addition, we show that MatchFixAgent can repair 50.6% of inequivalent translation, compared to prior work’s 18.5%.
    BibTeX
    @inproceedings{ibrahimzada2026matchfixagent,
      title = 	 {{M}atch{F}ix{A}gent: Language-Agnostic Autonomous Repository-Level Code Translation Validation and Repair},
      author =       {Ibrahimzada, Ali Reza and Paulsen, Brandon and Jabbarvand, Reyhaneh and Dodds, Joey and Kroening, Daniel},
      booktitle = 	 {Proceedings of the 43rd International Conference on Machine Learning},
      pages = 	 {49428--49451},
      year = 	 {2026},
      editor = 	 {Zhang, Tong and Dudik, Miroslav and Jaggi, Martin and Agarwal, Alekh and Li, Sharon and Schuurmans, Dale and Zhu, Jerry and Berkenkamp, Felix and Dong, Hanze and Bietti, Alberto},
      volume = 	 {306},
      series = 	 {Proceedings of Machine Learning Research},
      month = 	 {06--11 Jul},
      publisher =    {PMLR},
      pdf = 	 {https://raw.githubusercontent.com/mlresearch/v306/main/assets/ibrahimzada26a/ibrahimzada26a.pdf},
      url = 	 {https://proceedings.mlr.press/v306/ibrahimzada26a.html},
      abstract = 	 {Code translation transforms source code from one programming language (PL) to another. Validating the functional equivalence of translation and repairing, if necessary, are critical steps in code translation. Existing automated validation and repair approaches struggle to generalize to many PLs due to high engineering overhead, and they rely on existing and often inadequate test suites, which results in false claims of equivalence and ineffective translation repair. To bridge this gap, we develop MatchFixAgent, a large language model (LLM)-based, PL-agnostic framework for equivalence validation and repair of translations. MatchFixAgent features a multi-agent architecture that divides equivalence validation into several sub-tasks to ensure thorough and consistent semantic analysis of the translation. We compare MatchFixAgent’s validation and repair results with four repository-level code translation techniques. Our results demonstrate that MatchFixAgent produces (in)equivalence verdicts for $99.2$% of translation pairs, with the same equivalence validation result as prior work on $72.8$% of them. When MatchFixAgent’s result disagrees with prior work, we find that $60.7$% of the time MatchFixAgent’s result is actually correct. In addition, we show that MatchFixAgent can repair $50.6$% of inequivalent translation, compared to prior work’s $18.5$%.}
    }
  4. Nature

    Predicting student grades via adaptive multi-level learning models

    Ali Reza Ibrahimzada, Kerem Kosif, Ahmed Said Gulsen, Yavuz Selim Yaglica and Ali Cakmak
    Scientific Reports 0 citations 1 star
    Abstract
    Educational institutions increasingly rely on intelligent systems to extract actionable insights from student data. One critical application is the early prediction of student performance in specific courses, which can inform academic advising, course selection, and targeted interventions. This paper proposes an adaptive multi-level prediction framework that segments student-course data into homogeneous groups and assigns a temporally validated specialist model to each group. The framework is model-agnostic: it accepts any learner implementing standard fit/predict interfaces, including linear regression, ensemble methods, neural networks, and collaborative filtering. To combat temporal data sparsity common in volatile or block-cohort curriculum structures, the framework incorporates an automated data-density fallback guard that dynamically transitions isolated data slices to localized validation pools. Systematic validation on two large-scale, real-world higher education datasets demonstrates strong cross-institutional generalizability. On a dense traditional university dataset, the framework yields an 18.6% RMSE improvement over global baselines, with model selection converging heavily on tree-based ensembles. Conversely, on a volatile sparse course dataset, the pipeline automatically uncovers an extraordinarily diverse model ecosystem–selecting neural layers for 48% of clusters and triggering the collaborative filtering 22% of chronological windows. Backed by asymptotic significance testing (p-val. }}< 10^{-207}}}), these results prove that the framework effectively shifts the configuration burden from manual heuristics to self-correcting, data-driven optimization.
    BibTeX
    @Article{ibrahimzada2026predicting,
        author={Ibrahimzada, Ali Reza
        and Kosif, Kerem
        and Gulsen, Ahmed Said
        and Yaglica, Yavuz Selim
        and Cakmak, Ali},
        title={Predicting student grades via adaptive multi-level learning models},
        journal={Scientific Reports},
        year={2026},
        month={Jun},
        day={04},
        volume={16},
        number={1},
        pages={25619},
        abstract={Educational institutions increasingly rely on intelligent systems to extract actionable insights from student data. One critical application is the early prediction of student performance in specific courses, which can inform academic advising, course selection, and targeted interventions. This paper proposes an adaptive multi-level prediction framework that segments student-course data into homogeneous groups and assigns a temporally validated specialist model to each group. The framework is model-agnostic: it accepts any learner implementing standard fit/predict interfaces, including linear regression, ensemble methods, neural networks, and collaborative filtering. To combat temporal data sparsity common in volatile or block-cohort curriculum structures, the framework incorporates an automated data-density fallback guard that dynamically transitions isolated data slices to localized validation pools. Systematic validation on two large-scale, real-world higher education datasets demonstrates strong cross-institutional generalizability. On a dense traditional university dataset, the framework yields an 18.6{\%} RMSE improvement over global baselines, with model selection converging heavily on tree-based ensembles. Conversely, on a volatile sparse course dataset, the pipeline automatically uncovers an extraordinarily diverse model ecosystem--selecting neural layers for 48{\%} of clusters and triggering the collaborative filtering 22{\%} of chronological windows. Backed by asymptotic significance testing (p-val. {\$}{\$}< 10^{\{}-207{\}}{\$}{\$}), these results prove that the framework effectively shifts the configuration burden from manual heuristics to self-correcting, data-driven optimization.},
        issn={2045-2322},
        doi={10.1038/s41598-026-55750-z},
        url={https://doi.org/10.1038/s41598-026-55750-z}
    }

2025

  1. pre-print

    Advancing Automated In-Isolation Validation in Repository-Level Code Translation

    arXiv preprint arXiv:2511.21878 4 citations
    Abstract
    Repository-level code translation aims to migrate entire repositories across programming languages while preserving functionality automatically. Despite advancements in repository-level code translation, validating the translations remains challenging. This paper proposes TRAM, which combines context-aware type resolution with mock-based in-isolation validation to achieve high-quality translations between programming languages. Prior to translation, TRAM retrieves API documentation and contextual code information for each variable type in the source language. It then prompts a large language model (LLM) with retrieved contextual information to resolve type mappings across languages with precise semantic interpretations. Using the automatically constructed type mapping, TRAM employs a custom serialization/deserialization workflow that automatically constructs equivalent mock objects in the target language. This enables each method fragment to be validated in isolation, without the high cost of using agents for translation validation, or the heavy manual effort required by existing approaches that rely on language interoperability. TRAM demonstrates state-of-the-art performance in Java-to-Python translation, underscoring the effectiveness of its integration of RAG-based type resolution with reliable in-isolation validation.
    BibTeX
    @article{ke2025advancing,
        title={Advancing Automated In-Isolation Validation in Repository-Level Code Translation},
        author={Ke, Kaiyao and Ibrahimzada, Ali Reza and Pan, Rangeet and Sinha, Saurabh and Jabbarvand, Reyhaneh},
        journal={arXiv preprint arXiv:2511.21878},
        year={2025}
    }
  2. FSE’25
    ACM Artifacts Available ACM Artifacts Evaluated: Functional

    AlphaTrans: A Neuro-Symbolic Compositional Approach for Repository-Level Code Translation and Validation

    In Proceedings of the ACM Conference on Foundations of Software Engineering (FSE’25), June 23-27, 2025, Trondheim, Norway 104 citations 38 stars
    Abstract
    Code translation transforms programs from one programming language (PL) to another. One prominent use case is application modernization to enhance maintainability and reliability. Several rule-based transpilers have been designed to automate code translation between different pairs of PLs. However, the rules can become obsolete as the PLs evolve and cannot generalize to other PLs. Recent studies have explored the automation of code translation using Large Language Models (LLMs). One key observation is that such techniques may work well for crafted benchmarks but fail to generalize to the scale and complexity of real-world projects with inter- and intra-class dependencies, custom types, PL-specific features, etc. We propose AlphaTrans, a neuro-symbolic approach to automate repository-level code translation. AlphaTrans translates both source and test code, and employs multiple levels of validation to ensure the translation preserves the functionality of the source program. To break down the problem for LLMs, AlphaTrans leverages program analysis to decompose the program into fragments and translates them in the reverse call order.We leveraged AlphaTrans to translate ten real-world open-source projects consisting of ⟨836, 8575, 2719⟩ (application and test) classes, (application and test) methods, and unit tests. AlphaTrans breaks down these projects into 17874 fragments and translates the entire repository. 96.40% of the translated fragments are syntactically correct, and AlphaTrans validates the translations’ runtime behavior and functional correctness for 27.03% and 25.14% of the application method fragments. On average, integrated translation and validation takes 34 hours (min=3, max=121) to translate a project, showing its scalability in practice. For the syntactically or semantically incorrect translations, AlphaTrans generates a report including existing translation, stack trace, test errors, or assertion failures. We provided these artifacts to two developers to fix the translation bugs in four projects. They fixed the issues in 20.1 hours on average (5.5 hours for the smallest and 34 hours for the largest project) and achieved all passing tests. Without AlphaTrans, translating and validating such big projects could take weeks, if not months.
    BibTeX
    @article{ibrahimzada2025alphatrans,
        author = {Ibrahimzada, Ali Reza and Ke, Kaiyao and Pawagi, Mrigank and Abid, Muhammad Salman and Pan, Rangeet and Sinha, Saurabh and Jabbarvand, Reyhaneh},
        title = {AlphaTrans: A Neuro-Symbolic Compositional Approach for Repository-Level Code Translation and Validation},
        year = {2025},
        issue_date = {July 2025},
        publisher = {Association for Computing Machinery},
        address = {New York, NY, USA},
        volume = {2},
        number = {FSE},
        url = {https://doi.org/10.1145/3729379},
        doi = {10.1145/3729379},
        abstract = {Code translation transforms programs from one programming language (PL) to another. One prominent use case is application modernization to enhance maintainability and reliability. Several rule-based transpilers have been designed to automate code translation between different pairs of PLs. However, the rules can become obsolete as the PLs evolve and cannot generalize to other PLs. Recent studies have explored the automation of code translation using Large Language Models (LLMs). One key observation is that such techniques may work well for crafted benchmarks but fail to generalize to the scale and complexity of real-world projects with inter- and intra-class dependencies, custom types, PL-specific features, etc. We propose AlphaTrans, a neuro-symbolic approach to automate repository-level code translation. AlphaTrans translates both source and test code, and employs multiple levels of validation to ensure the translation preserves the functionality of the source program. To break down the problem for LLMs, AlphaTrans leverages program analysis to decompose the program into fragments and translates them in the reverse call order.We leveraged AlphaTrans to translate ten real-world open-source projects consisting of ⟨836, 8575, 2719⟩ (application and test) classes, (application and test) methods, and unit tests. AlphaTrans breaks down these projects into 17874 fragments and translates the entire repository. 96.40\% of the translated fragments are syntactically correct, and AlphaTrans validates the translations’ runtime behavior and functional correctness for 27.03\% and 25.14\% of the application method fragments. On average, integrated translation and validation takes 34 hours (min=3, max=121) to translate a project, showing its scalability in practice. For the syntactically or semantically incorrect translations, AlphaTrans generates a report including existing translation, stack trace, test errors, or assertion failures. We provided these artifacts to two developers to fix the translation bugs in four projects. They fixed the issues in 20.1 hours on average (5.5 hours for the smallest and 34 hours for the largest project) and achieved all passing tests. Without AlphaTrans, translating and validating such big projects could take weeks, if not months.},
        journal = {Proc. ACM Softw. Eng.},
        month = jun,
        articleno = {FSE109},
        numpages = {23},
        keywords = {Neuro-Symbolic Code Translation and Validation}
    }
  3. SCAM’25

    Challenging Bug Prediction and Repair Models with Synthetic Bugs

    Ali Reza Ibrahimzada, Yang Chen, Ryan Rong and Reyhaneh Jabbarvand
    In Proceedings of the 25th IEEE International Conference on Source Code Analysis & Manipulation (SCAM’25), September 7-12, 2025, Auckland, New Zealand 30 citations 6 stars
    Abstract
    Bugs are essential in software engineering; many research studies in the past decades have been proposed to detect, localize, and repair bugs in software systems. Effectiveness evaluation of such techniques requires complex bugs, i.e., those that are hard to detect through testing and hard to repair through debugging. From the classic software engineering point of view, a hard-to-repair bug differs from the correct code in multiple locations, making it hard to localize and repair. Hard-to-detect bugs, on the other hand, manifest themselves under specific test inputs and reachability conditions. These two objectives, i.e., generating hard-to-detect and hard-to-repair bugs, are mostly aligned; a bug generation technique can change multiple statements to be covered only under a specific set of inputs. However, these two objectives conflict in the learning-based techniques: A bug should have a similar code representation to the correct code in the training data to challenge a bug prediction model to distinguish them. The hard-to-repair bug definition remains the same but with a caveat: the more a bug differs from the original code (at multiple locations), the more distant their representations are and easier to detect. This demands new techniques to generate bugs to complement existing bug datasets to challenge learning-based bug prediction and repair techniques.(p)(/p)We propose BugFarm to transform arbitrary code into multiple hard-to-detect and hard-to-repair bugs. BugFarm mutates code in multiple locations (hard-to-repair) but leverages attention analysis to only change the least attended locations by the underlying model (hard-to-detect). Our comprehensive evaluation of 435k+ bugs from over 1.9M mutants generated by BugFarm and two alternative approaches demonstrates our superiority in generating bugs that are hard to detect by learning-based bug prediction approaches (up to 40.53% higher False Negative Rate and 10.76%, 5.2%, 28.93%, and 20.53% lower Accuracy, Precision, Recall, and F1 score) and hard to repair by state-of-the-art learning-based program repair technique (28% repair success rate compared to 36% and 49% of LEAM and μBERT bugs). BugFarm is efficient, i.e., it takes nine seconds to mutate a code with no training overhead.
    BibTeX
    @INPROCEEDINGS{ibrahimzada2025challenging,
        author={Ibrahimzada, Ali Reza and Chen, Yang and Rong, Ryan and Jabbarvand, Reyhaneh},
        booktitle={2025 IEEE International Conference on Source Code Analysis & Manipulation (SCAM)}, 
        title={Challenging Bug Prediction and Repair Models with Synthetic Bugs}, 
        year={2025},
        volume={},
        number={},
        pages={133-144},
        keywords={Training;Codes;Computer bugs;Training data;Transforms;Maintenance engineering;Predictive models;Software systems;Software engineering;Testing;Bug Generation;Bug Prediction;Interpretation},
        doi={10.1109/SCAM67354.2025.00021}
    }

2024

  1. MS THESIS

    Bridging the Gap between Testing and Debugging through Explainable Deep Oracles

    Ali Reza Ibrahimzada
    MS Thesis, University of Illinois Urbana-Champaign 0 citations 11 stars
    Abstract
    Automation of test oracles is one of the most challenging facets of software testing, but remains comparatively less addressed compared to automated test input generation. Test oracles rely on a ground-truth that can distinguish between the correct and buggy behavior to determine whether a test fails (detects a bug) or passes. What makes the oracle problem challenging and undecidable is the assumption that the ground-truth should know the exact expected, correct, or buggy behavior. However, we argue that one can still build an accurate oracle without knowing the exact correct or buggy behavior, but how these two might differ. This paper presents , a learning-based approach that in the absence of test assertions or other types of oracle, can determine whether a unit test passes or fails on a given method under test (MUT). To build the ground-truth, jointly embeds unit tests and the implementation of MUTs into a unified vector space, in such a way that the neural representation of tests are similar to that of MUTs they pass on them, but dissimilar to MUTs they fail on them. The classifier built on top of this vector representation serves as the oracle to generate “fail” labels, when test inputs detect a bug in MUT or “pass” labels, otherwise. Our extensive experiments on applying to more than 5K unit tests from a diverse set of open-source Java projects show that the produced oracle is (1) effective in predicting the fail or pass labels, achieving an overall accuracy, precision, recall, and F1 measure of 93%, 86%, 94%, and 90%, (2) generalizable, predicting the labels for the unit test of projects that were not in training or validation set with negligible performance drop, and (3) efficient, detecting the existence of bugs in only 6.5 milliseconds on average. Moreover, by interpreting the neural model and looking at it beyond a closed-box solution, we confirm that the oracle is valid, i.e., it predicts the labels through learning relevant features.
    BibTeX
    @phdthesis{ibrahimzada2024bridging,
      title={Bridging the gap between testing and debugging through explainable deep oracles},
      author={Ibrahimzada, Ali Reza},
      year={2024},
      school={University of Illinois at Urbana-Champaign}
    }
  2. ICSE’24
    ACM Artifacts Available ACM Artifacts Evaluated: Reusable

    Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code

    Rangeet Pan*, Ali Reza Ibrahimzada*, Rahul Krishna, Divya Sankar, Lambert Pouguem Wassi, Michele Merler, Boris Sobolev, Raju Pavuluri, Saurabh Sinha and Reyhaneh Jabbarvand
    (* equal contribution)
    In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering (ICSE ’24), April 14-20, 2024, Lisbon, Portugal 457 citations 54 stars
    Abstract
    Code translation aims to convert source code from one programming language (PL) to another. Given the promising abilities of large language models (LLMs) in code synthesis, researchers are exploring their potential to automate code translation. The prerequisite for advancing the state of LLM-based code translation is to understand their promises and limitations over existing techniques. To that end, we present a large-scale empirical study to investigate the ability of general LLMs and code LLMs for code translation across pairs of different languages, including C, C++, Go, Java, and Python. Our study, which involves the translation of 1,700 code samples from three benchmarks and two real-world projects, reveals that LLMs are yet to be reliably used to automate code translation—with correct translations ranging from 2.1% to 47.3% for the studied LLMs. Further manual investigation of unsuccessful translations identifies 15 categories of translation bugs. We also compare LLM-based code translation with traditional non-LLM-based approaches. Our analysis shows that these two classes of techniques have their own strengths and weaknesses. Finally, insights from our study suggest that providing more context to LLMs during translation can help them produce better results. To that end, we propose a prompt-crafting approach based on the symptoms of erroneous translations; this improves the performance of LLM-based code translation by 5.5% on average. Our study is the first of its kind, in terms of scale and breadth, that provides insights into the current limitations of LLMs in code translation and opportunities for improving them. Our dataset—consisting of 1,700 code samples in five PLs with 10K+ tests, 43K+ translated code, 1,748 manually labeled bugs, and 1,365 bug-fix pairs—can help drive research in this area.
    BibTeX
    @inproceedings{pan2024lost,
        author = {Pan, Rangeet and Ibrahimzada, Ali Reza and Krishna, Rahul and Sankar, Divya and Wassi, Lambert Pouguem and Merler, Michele and Sobolev, Boris and Pavuluri, Raju and Sinha, Saurabh and Jabbarvand, Reyhaneh},
        title = {Lost in Translation: A Study of Bugs Introduced by Large Language Models while Translating Code},
        year = {2024},
        isbn = {9798400702174},
        publisher = {Association for Computing Machinery},
        address = {New York, NY, USA},
        url = {https://doi.org/10.1145/3597503.3639226},
        doi = {10.1145/3597503.3639226},
        abstract = {Code translation aims to convert source code from one programming language (PL) to another. Given the promising abilities of large language models (LLMs) in code synthesis, researchers are exploring their potential to automate code translation. The prerequisite for advancing the state of LLM-based code translation is to understand their promises and limitations over existing techniques. To that end, we present a large-scale empirical study to investigate the ability of general LLMs and code LLMs for code translation across pairs of different languages, including C, C++, Go, Java, and Python. Our study, which involves the translation of 1,700 code samples from three benchmarks and two real-world projects, reveals that LLMs are yet to be reliably used to automate code translation---with correct translations ranging from 2.1\% to 47.3\% for the studied LLMs. Further manual investigation of unsuccessful translations identifies 15 categories of translation bugs. We also compare LLM-based code translation with traditional non-LLM-based approaches. Our analysis shows that these two classes of techniques have their own strengths and weaknesses. Finally, insights from our study suggest that providing more context to LLMs during translation can help them produce better results. To that end, we propose a prompt-crafting approach based on the symptoms of erroneous translations; this improves the performance of LLM-based code translation by 5.5\% on average. Our study is the first of its kind, in terms of scale and breadth, that provides insights into the current limitations of LLMs in code translation and opportunities for improving them. Our dataset---consisting of 1,700 code samples in five PLs with 10K+ tests, 43K+ translated code, 1,748 manually labeled bugs, and 1,365 bug-fix pairs---can help drive research in this area.},
        booktitle = {Proceedings of the IEEE/ACM 46th International Conference on Software Engineering},
        articleno = {82},
        numpages = {13},
        keywords = {code translation, bug taxonomy, llm},
        location = {Lisbon, Portugal},
        series = {ICSE '24}
    }
  3. ICSE’24 SRC

    Program Decomposition and Translation with Static Analysis

    Ali Reza Ibrahimzada
    In Proceedings of the 46th IEEE/ACM International Conference on Software Engineering Companion (ICSE ’24 Companion), April 14-20, 2024, Lisbon, Portugal 11 citations 🏆 Ranked 3rd in IEEE/ACM Student Research Competition at ICSE’24
    Abstract
    The rising popularity of Large Language Models (LLMs) has motivated exploring their use in code-related tasks. Code LLMs with more than millions of parameters are trained on a massive amount of code in different Programming Languages (PLs). Such models are used for automating various Software Engineering (SE) tasks using prompt engineering. However, given the very large size of industry-scale project files, a major issue of these LLMs is their limited context window size, motivating the question of "Can these LLMs process very large files and can we effectively perform prompt engineering?". Code translation aims to convert source code from one PL to another. In this work, we assess the effect of method-level program decomposition on context window of LLMs and investigate how this approach can enable translation of very large files which originally could not be done due to out-of-context issue. Our observations from 20 well-known java projects and approximately 60K methods suggest that method-level program decomposition significantly improves the limited context window problem of LLMs by 99.5%. Furthermore, our empirical analysis indicate that with method-level decomposition, each input fragment on average only consumes 5% of the context window, leaving more context space for prompt engineering and the output. Finally, we investigate the effectiveness of a Call Graph (CG) approach for translating very large files when doing method-level program decomposition.
    BibTeX
    @inproceedings{ibrahimzada2024program,
        author = {Ibrahimzada, Ali Reza},
        title = {Program Decomposition and Translation with Static Analysis},
        year = {2024},
        isbn = {9798400705021},
        publisher = {Association for Computing Machinery},
        address = {New York, NY, USA},
        url = {https://doi.org/10.1145/3639478.3641226},
        doi = {10.1145/3639478.3641226},
        abstract = {The rising popularity of Large Language Models (LLMs) has motivated exploring their use in code-related tasks. Code LLMs with more than millions of parameters are trained on a massive amount of code in different Programming Languages (PLs). Such models are used for automating various Software Engineering (SE) tasks using prompt engineering. However, given the very large size of industry-scale project files, a major issue of these LLMs is their limited context window size, motivating the question of "Can these LLMs process very large files and can we effectively perform prompt engineering?". Code translation aims to convert source code from one PL to another. In this work, we assess the effect of method-level program decomposition on context window of LLMs and investigate how this approach can enable translation of very large files which originally could not be done due to out-of-context issue. Our observations from 20 well-known java projects and approximately 60K methods suggest that method-level program decomposition significantly improves the limited context window problem of LLMs by 99.5\%. Furthermore, our empirical analysis indicate that with method-level decomposition, each input fragment on average only consumes 5\% of the context window, leaving more context space for prompt engineering and the output. Finally, we investigate the effectiveness of a Call Graph (CG) approach for translating very large files when doing method-level program decomposition.},
        booktitle = {Proceedings of the 2024 IEEE/ACM 46th International Conference on Software Engineering: Companion Proceedings},
        pages = {453–455},
        numpages = {3},
        location = {Lisbon, Portugal},
        series = {ICSE-Companion '24}
    }
  4. pre-print

    CodeMind: A Framework to Challenge Large Language Models for Code Reasoning

    Changshu Liu, Shizhuo Dylan Zhang, Ali Reza Ibrahimzada and Reyhaneh Jabbarvand
    arXiv preprint arXiv:2402.09664 99 citations 44 stars
    Abstract
    Solely relying on test passing to evaluate Large Language Models (LLMs) for code synthesis may result in unfair assessment or promoting models with data leakage. As an alternative, we introduce CodeMind, a framework designed to gauge the code reasoning abilities of LLMs. CodeMind currently supports three code reasoning tasks: Independent Execution Reasoning (IER), Dependent Execution Reasoning (DER), and Specification Reasoning (SR). The first two evaluate models to predict the execution output of an arbitrary code or code the model could correctly synthesize. The third one evaluates the extent to which LLMs implement the specified expected behavior. Our extensive evaluation of nine LLMs across five benchmarks in two different programming languages using CodeMind shows that LLMs fairly follow control flow constructs and, in general, explain how inputs evolve to output, specifically for simple programs and the ones they can correctly synthesize. However, their performance drops for code with higher complexity, non-trivial logical and arithmetic operators, non-primitive types, and API calls. Furthermore, we observe that, while correlated, specification reasoning (essential for code synthesis) does not imply execution reasoning (essential for broader programming tasks such as testing and debugging): ranking LLMs based on test passing can be different compared to code reasoning.
    BibTeX
    @article{liu2024codemind,
      title={Codemind: A framework to challenge large language models for code reasoning},
      author={Liu, Changshu and Zhang, Shizhuo Dylan and Ibrahimzada, Ali Reza and Jabbarvand, Reyhaneh},
      journal={arXiv preprint arXiv:2402.09664},
      year={2024}
    }

2023

  1. MBEC

    Predicting the predisposition to colorectal cancer based on SNP profiles of immune phenotypes using supervised learning models

    Ali Cakmak, Huzeyfe Ayaz, Soykan Arıkan, Ali Reza Ibrahimzada, Şeyda Demirkol, Dilara Sönmez, Mehmet T. Hakan, Saime T. Sürmen, Cem Horozoğlu, Mehmet B. Doğan, Özlem Küçükhüseyin, Canan Cacına, Bayram Kıran, Ümit Zeybek, Mehmet Baysan and İlhan Yaylım
    Medical & Biological Engineering & Computing, Springer Berlin Heidelberg, Vol. 61, 243–258, 2023 6 citations 0 stars
    Abstract
    This study explores the machine learning-based assessment of predisposition to colorectal cancer based on single nucleotide polymorphisms (SNP). Such a computational approach may be used as a risk indicator and an auxiliary diagnosis method that complements the traditional methods such as biopsy and CT scan. Moreover, it may be used to develop a low-cost screening test for the early detection of colorectal cancers to improve public health. We employ several supervised classification algorithms. Besides, we apply data imputation to fill in the missing genotype values. The employed dataset includes SNPs observed in particular colorectal cancer-associated genomic loci that are located within DNA regions of 11 selected genes obtained from 115 individuals. We make the following observations: (i) random forest-based classifier using one-hot encoding and K-nearest neighbor (KNN)-based imputation performs the best among the studied classifiers with an F1 score of 89% and area under the curve (AUC) score of 0.96. (ii) One-hot encoding together with K-nearest neighbor-based data imputation increases the F1 scores by around 26% in comparison to the baseline approach which does not employ them. (iii) The proposed model outperforms a commonly employed state-of-the-art approach, ColonFlag, under all evaluated settings by up to 24% in terms of the AUC score. Based on the high accuracy of the constructed predictive models, the studied 11 genes may be considered a gene panel candidate for colon cancer risk screening.
    BibTeX
    @Article{cakmak2023predictingthe,
        author={Cakmak, Ali
        and Ayaz, Huzeyfe
        and Ar{\i}kan, Soykan
        and Ibrahimzada, Ali R.
        and Demirkol, {\c{S}}eyda
        and S{\"o}nmez, Dilara
        and Hakan, Mehmet T.
        and S{\"u}rmen, Saime T.
        and Horozo{\u{g}}lu, Cem
        and Do{\u{g}}an, Mehmet B.
        and K{\"u}{\c{c}}{\"u}kh{\"u}seyin, {\"O}zlem
        and Cac{\i}na, Canan
        and K{\i}ran, Bayram
        and Zeybek, {\"U}mit
        and Baysan, Mehmet
        and Yayl{\i}m, {\.{I}}lhan},
        title={Predicting the predisposition to colorectal cancer based on SNP profiles of immune phenotypes using supervised learning models},
        journal={Medical {\&} Biological Engineering {\&} Computing},
        year={2023},
        month={Jan},
        day={01},
        volume={61},
        number={1},
        pages={243-258},
        abstract={This study explores the machine learning-based assessment of predisposition to colorectal cancer based on single nucleotide polymorphisms (SNP). Such a computational approach may be used as a risk indicator and an auxiliary diagnosis method that complements the traditional methods such as biopsy and CT scan. Moreover, it may be used to develop a low-cost screening test for the early detection of colorectal cancers to improve public health. We employ several supervised classification algorithms. Besides, we apply data imputation to fill in the missing genotype values. The employed dataset includes SNPs observed in particular colorectal cancer-associated genomic loci that are located within DNA regions of 11 selected genes obtained from 115 individuals. We make the following observations: (i) random forest-based classifier using one-hot encoding and K-nearest neighbor (KNN)-based imputation performs the best among the studied classifiers with an F1 score of 89{\%} and area under the curve (AUC) score of 0.96. (ii) One-hot encoding together with K-nearest neighbor-based data imputation increases the F1 scores by around 26{\%} in comparison to the baseline approach which does not employ them. (iii) The proposed model outperforms a commonly employed state-of-the-art approach, ColonFlag, under all evaluated settings by up to 24{\%} in terms of the AUC score. Based on the high accuracy of the constructed predictive models, the studied 11 genes may be considered a gene panel candidate for colon cancer risk screening.},
        issn={1741-0444},
        doi={10.1007/s11517-022-02707-9},
        url={https://doi.org/10.1007/s11517-022-02707-9}
    }

2022

  1. ESEC/FSE’22
    ACM Artifacts Available ACM Artifacts Evaluated: Functional

    Perfect Is the Enemy of Test Oracle

    Ali Reza Ibrahimzada, Yigit Varli, Dilara Tekinoglu and Reyhaneh Jabbarvand
    In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE ’22), November 14–18, 2022, Singapore, Singapore 49 citations 11 stars
    Abstract
    Automation of test oracles is one of the most challenging facets of software testing, but remains comparatively less addressed compared to automated test input generation. Test oracles rely on a ground-truth that can distinguish between the correct and buggy behavior to determine whether a test fails (detects a bug) or passes. What makes the oracle problem challenging and undecidable is the assumption that the ground-truth should know the exact expected, correct, or buggy behavior. However, we argue that one can still build an accurate oracle without knowing the exact correct or buggy behavior, but how these two might differ. This paper presents , a learning-based approach that in the absence of test assertions or other types of oracle, can determine whether a unit test passes or fails on a given method under test (MUT). To build the ground-truth, jointly embeds unit tests and the implementation of MUTs into a unified vector space, in such a way that the neural representation of tests are similar to that of MUTs they pass on them, but dissimilar to MUTs they fail on them. The classifier built on top of this vector representation serves as the oracle to generate “fail” labels, when test inputs detect a bug in MUT or “pass” labels, otherwise. Our extensive experiments on applying to more than 5K unit tests from a diverse set of open-source Java projects show that the produced oracle is (1) effective in predicting the fail or pass labels, achieving an overall accuracy, precision, recall, and F1 measure of 93%, 86%, 94%, and 90%, (2) generalizable, predicting the labels for the unit test of projects that were not in training or validation set with negligible performance drop, and (3) efficient, detecting the existence of bugs in only 6.5 milliseconds on average. Moreover, by interpreting the neural model and looking at it beyond a closed-box solution, we confirm that the oracle is valid, i.e., it predicts the labels through learning relevant features.
    BibTeX
    @inproceedings{ibrahimzada2022perfect,
        author = {Ibrahimzada, Ali Reza and Varli, Yigit and Tekinoglu, Dilara and Jabbarvand, Reyhaneh},
        title = {Perfect is the enemy of test oracle},
        year = {2022},
        isbn = {9781450394130},
        publisher = {Association for Computing Machinery},
        address = {New York, NY, USA},
        url = {https://doi.org/10.1145/3540250.3549086},
        doi = {10.1145/3540250.3549086},
        abstract = {Automation of test oracles is one of the most challenging facets of software testing, but remains comparatively less addressed compared to automated test input generation. Test oracles rely on a ground-truth that can distinguish between the correct and buggy behavior to determine whether a test fails (detects a bug) or passes. What makes the oracle problem challenging and undecidable is the assumption that the ground-truth should know the exact expected, correct, or buggy behavior. However, we argue that one can still build an accurate oracle without knowing the exact correct or buggy behavior, but how these two might differ. This paper presents , a learning-based approach that in the absence of test assertions or other types of oracle, can determine whether a unit test passes or fails on a given method under test (MUT). To build the ground-truth, jointly embeds unit tests and the implementation of MUTs into a unified vector space, in such a way that the neural representation of tests are similar to that of MUTs they pass on them, but dissimilar to MUTs they fail on them. The classifier built on top of this vector representation serves as the oracle to generate “fail” labels, when test inputs detect a bug in MUT or “pass” labels, otherwise. Our extensive experiments on applying to more than 5K unit tests from a diverse set of open-source Java projects show that the produced oracle is (1) effective in predicting the fail or pass labels, achieving an overall accuracy, precision, recall, and F1 measure of 93\%, 86\%, 94\%, and 90\%, (2) generalizable, predicting the labels for the unit test of projects that were not in training or validation set with negligible performance drop, and (3) efficient, detecting the existence of bugs in only 6.5 milliseconds on average. Moreover, by interpreting the neural model and looking at it beyond a closed-box solution, we confirm that the oracle is valid, i.e., it predicts the labels through learning relevant features.},
        booktitle = {Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering},
        pages = {70–81},
        numpages = {12},
        keywords = {Deep Learning, Software Testing, Test Automation, Test Oracle},
        location = {Singapore, Singapore},
        series = {ESEC/FSE 2022}
    }

2020

  1. JBCB

    Scalable classification of organisms into a taxonomy using hierarchical supervised learners

    Gihad N. Sohsah, Ali Reza Ibrahimzada, Huzeyfe Ayaz and Ali Cakmak
    Journal of Bioinformatics and Computational Biology, World Scientific Publishing Co., Vol. 18, No. 05, 2020 7 citations 1 star
    Abstract
    Accurately identifying organisms based on their partially available genetic material is an important task to explore the phylogenetic diversity in an environment. Specific fragments in the DNA sequence of a living organism have been defined as DNA barcodes and can be used as markers to identify species efficiently and effectively. The existing DNA barcode-based classification approaches suffer from three major issues: (i) most of them assume that the classification is done within a given taxonomic class and/or input sequences are pre-aligned, (ii) highly performing classifiers, such as SVM, cannot scale to large taxonomies due to high memory requirements, (iii) mutations and noise in input DNA sequences greatly reduce the taxonomic classification score. In order to address these issues, we propose a multi-level hierarchical classifier framework to automatically assign taxonomy labels to DNA sequences. We utilize an alignment-free approach called spectrum kernel method for feature extraction. We build a proof-of-concept hierarchical classifier with two levels, and evaluated it on real DNA sequence data from barcode of life data systems. We demonstrate that the proposed framework provides higher f1-score than regular classifiers. Besides, hierarchical framework scales better to large datasets enabling researchers to employ classifiers with high classification performance and high memory requirement on large datasets. Furthermore, we show that the proposed framework is more robust to mutations and noise in sequence data than the non-hierarchical classifiers.
    BibTeX
    @article{sohsah2020scalable,
        author = {Sohsah, Gihad N. and Ibrahimzada, Ali Reza and Ayaz, Huzeyfe and Cakmak, Ali},
        title = {Scalable classification of organisms into a taxonomy using hierarchical supervised learners},
        journal = {Journal of Bioinformatics and Computational Biology},
        volume = {18},
        number = {05},
        pages = {2050026},
        year = {2020},
        doi = {10.1142/S0219720020500262},
        note ={PMID: 33125294},
        URL = {https://doi.org/10.1142/S0219720020500262},
        eprint = {https://doi.org/10.1142/S0219720020500262},
        abstract = {Accurately identifying organisms based on their partially available genetic material is an important task to explore the phylogenetic diversity in an environment. Specific fragments in the DNA sequence of a living organism have been defined as DNA barcodes and can be used as markers to identify species efficiently and effectively. The existing DNA barcode-based classification approaches suffer from three major issues: (i) most of them assume that the classification is done within a given taxonomic class and/or input sequences are pre-aligned, (ii) highly performing classifiers, such as SVM, cannot scale to large taxonomies due to high memory requirements, (iii) mutations and noise in input DNA sequences greatly reduce the taxonomic classification score. In order to address these issues, we propose a multi-level hierarchical classifier framework to automatically assign taxonomy labels to DNA sequences. We utilize an alignment-free approach called spectrum kernel method for feature extraction. We build a proof-of-concept hierarchical classifier with two levels, and evaluated it on real DNA sequence data from barcode of life data systems. We demonstrate that the proposed framework provides higher f1-score than regular classifiers. Besides, hierarchical framework scales better to large datasets enabling researchers to employ classifiers with high classification performance and high memory requirement on large datasets. Furthermore, we show that the proposed framework is more robust to mutations and noise in sequence data than the non-hierarchical classifiers.}
    }

2019

  1. IMMC

    Predicting the Predisposition to Colorectal Cancer based on SNP Profiles of Immune Checkpoints Using Supervised Learning Models

    Ali Cakmak, Ali Reza Ibrahimzada, Soykan Arikan, Huzeyfe Ayaz, Seyda Demirkol, Dilara Sonmez, Mehmet Tolgahan Hakan, Saime Turan Surmen, Cem Horozoglu, Ozlem Kucukhuseyin, Canan Cacina, Bayram Kiran, Umit Zeybek, Mehmet Baysan and Ilhan Yaylim
    VII. International Molecular Medicine Congress, September 5-7, 2019, Istanbul, Turkey 0 citations
    Abstract
    The goal this analysis is to explore the machine learning-based assessment of predisposition to colorectal cancer based on the single nucleotide polymorphisms (SNP). Such a computational approach may be used as a risk indicator and an auxiliary diagnosis method complementary to other diagnosis tools, such as, biopsy, CT scan, and MRI. Moreover, it may be used to develop a low-cost screening test for early detection of colorectal cancers to improve the public health.