ember

16References

  1. Cloudflare. Clef-Flash. Model repository, Hugging Face Hub. https://huggingface.co/Cloudflare/clef-flash
  2. Model Context Protocol. Specification. https://modelcontextprotocol.io/specification
  3. Naeini, M. P., Cooper, G. F., and Hauskrecht, M. Obtaining well calibrated probabilities using Bayesian binning. AAAI, 2015. https://ojs.aaai.org/index.php/AAAI/article/view/9602
  4. Guo, C., Pleiss, G., Sun, Y., and Weinberger, K. Q. On calibration of modern neural networks. ICML, 2017. https://arxiv.org/abs/1706.04599
  5. Nixon, J., Dusenberry, M., Zhang, L., Jerfel, G., and Tran, D. Measuring calibration in deep learning. CVPR Workshops, 2019. https://arxiv.org/abs/1904.01685
  6. Brier, G. W. Verification of forecasts expressed in terms of probability. Monthly Weather Review 78(1), 1950. https://doi.org/10.1175/1520-0493(1950)078%3C0001:VOFEIT%3E2.0.CO;2
  7. Epstein, E. S. A scoring system for probability forecasts of ranked categories. Journal of Applied Meteorology 8(6), 1969. https://doi.org/10.1175/1520-0450(1969)008%3C0985:ASSFPF%3E2.0.CO;2
  8. Opitz, J., and Burst, S. Macro F1 and Macro F1. arXiv:1911.03347, 2019. https://arxiv.org/abs/1911.03347
  9. Efron, B., and Tibshirani, R. J. An Introduction to the Bootstrap. Chapman & Hall, 1993. https://doi.org/10.1201/9780429246593
  10. Geifman, Y., and El-Yaniv, R. Selective classification for deep neural networks. NeurIPS, 2017. https://arxiv.org/abs/1705.08500
  11. DeGroot, M. H., and Fienberg, S. E. The comparison and evaluation of forecasters. The Statistician 32(1-2), 1983. https://doi.org/10.2307/2987588
  12. Gebru, T., et al. Datasheets for datasets. Communications of the ACM 64(12), 2021. https://arxiv.org/abs/1803.09010